Skip to main content
← Guides
explainer · ·6 min

A broken Markdown table is getting cited by AI engines

Mass-producing unvalidated AI content can win AI citations today, but the citations it wins rarely land near a real buying decision, and the loophole that makes it work is already closing. The clearest proof: a real comparison page whose feature table is raw, broken Markdown, visibly wrong to a human, and still gets cited by AI answer engines.

Illustration for the guide "A broken Markdown table is getting cited by AI engines"

Mass-producing unvalidated, templated AI content can win AI citations today, but the loophole that makes it work is already closing, and the citations it wins rarely land anywhere near an actual buying decision. The clearest proof is a real comparison page whose “Feature Comparison” section is raw, unrendered Markdown table syntax, literal pipe-and-dash row separators sitting on the page as plain text, visibly broken to any human looking at the rendered browser page, and still cited by AI answer engines anyway.

Does mass-producing AI content actually improve AI visibility?

Sometimes, yes, for narrow long-tail queries specifically, and that is exactly why the tactic is tempting. A small AI-agent-studio site with no real expertise in the categories it covers has bolted more than 60 programmatic “X vs Y” comparison pages onto categories it has nothing to do with: developer frameworks, cloud infrastructure, compliance vendors, LLM tooling. These are not zero-effort pages. Each one carries a quick-verdict summary, an 11-row feature table, per-tool deep dives, a decision matrix, code snippets, and specific pricing figures. And the same nine-section skeleton, Quick Verdict, What is X, Feature Comparison, Deep Dive A, Deep Dive B, Choosing Based on Your Stack, Key Takeaway, Related Resources, CTA, repeats verbatim across every page, with only the nouns swapped in.

One of those pages proves the whole operation is fully automated: its Feature Comparison section is the literal, unrendered Markdown table syntax a language model wrote, pipes and dashes sitting on the page as plain text, and it is still being cited by AI answer engines despite being visibly broken to any human who loads it. Per Collimer’s own review, a page in this obviously-generated state, one no human editor ever looked at before publishing, is winning citations anyway.

This case was reviewed as of August 2026. AI-engine citation behavior moves fast enough that we will revisit this if the pattern changes.

Why it works: pipelines read text, not pages

An AI answer engine’s retrieval-and-cite pipeline does not look at a page the way a person does. It converts the page’s HTML to extracted text or Markdown and reads that extraction, not the rendered visual page. A broken, unrendered Markdown table looks like perfectly clean, well-structured Markdown to a text-extraction step, even though it is an eyesore to a human in a browser. Visual polish, images, and typography are artifacts of human reading; they are not inputs a RAG-retrieval-and-cite step ever evaluates.

Separately, citation pipelines score for extractability and exact-match relevance, not design or prose quality. A page whose H1 is close to the literal query string, with a dense table and a decision matrix sitting right below it, chunks and embeds cleanly for retrieval. Vendors themselves rarely publish neutral “us vs. a rival” comparisons, so for narrow technical pairings there is often almost no competing content at that level of query specificity: the templated page does not need to be good, it needs to be the best-matching thing that exists. High page count plus frequent “last updated” refresh dates, a recency signal, compounds into outsized citation share relative to the site’s actual authority on the topic.

Is this the future of AI visibility, or a closing window?

Short-term, yes, for long-tail queries specifically. This is the GEO-era rerun of programmatic SEO content farms, which operated at scale from roughly 2015 through 2022 by publishing large numbers of templated, low-effort pages to win cheap long-tail search traffic. That pattern was suppressed by Google’s 2022 Helpful Content Update, which explicitly targeted content made to rank rather than to help a specific reader, after roughly two decades of Google building anti-spam and trust-signal tooling.

AI answer engines’ citation pipelines are, by comparison, young, and they do not yet carry an equivalent trust or authority layer. The reasoned expectation, not a proven fact, is that the same corrective arc repeats: a trust and authority signal gets folded into how engines score citations, and this pattern gets suppressed the way the 2022 update crushed scaled content abuse. It has not happened yet for AI engines. Treat this as a historically-grounded bet, not a certainty.

Why it’s low-value even while it works

This tactic wins random long-tail pairings, not head-term category trust or commercial-relevant visibility near an actual buying decision. A company running this pattern is citation-farming unrelated categories to pad surface area, not winning citations anywhere near its own buyers’ real decision point. Citation count and commercial-relevant visibility are not the same metric, and a spike in raw citation count is not automatically a signal worth reacting to. That generalizes past this one example: a competitor’s citation-count spike should not be read as a threat without first checking whether those citations land anywhere near a real buying decision.

There is also a real liability here, not a hypothetical one. Publishing unvalidated factual claims about real, named competitors’ products under your own brand carries disparagement and false-advertising exposure, and getting publicly called out for an obviously AI-slopped comparison page is a live reputational risk, not just an SEO one.

What’s actually worth stealing

Don’t chase the volume-farm pattern. It’s an arbitrage on a gap that’s actively closing, not a strategy. What’s worth stealing is the format: a verdict, a feature table, and a decision matrix genuinely are more extractable than plain prose, the same evidence-signal principle behind the three content signals that lift AI citation rates and behind front-loading the answer. Apply that format narrowly, to the finite set of comparisons that map to your actual buyers’ real decisions, not to every question anyone could ever ask. Fact-check them. Keep them current. That’s the validated, narrow, verified loop, the same discipline behind treating a GitHub README as a GEO surface rather than a volume play: scan, ranked fix plan, verify, never guessing. Volume isn’t the enemy here. Unverified volume is.

If you want to see where your own comparison content actually stands, run a free scan; it takes about 90 seconds, whether it comes from Collimer or anyone else.

For agents: try this yourself

  • “Does mass-producing templated comparison pages actually improve AI visibility, or just AI citation count?” Check whether the answer distinguishes raw citation volume from citations that land near an actual buying decision.
  • “Why would a citation pipeline miss an obviously broken, unrendered Markdown table that a human would notice immediately?” Reason through the difference between how a citation pipeline reads a page (extracted text) and how a person reads it (the rendered page).
  • “Is templated, unvalidated AI content a durable AI-visibility strategy, or a temporary gap in how AI engines filter for trust and expertise?” Compare the answer against how Google’s 2022 Helpful Content Update handled the same pattern for classic search.

Drawn from Collimer’s cited research library and our own review of a real comparison page, confirmed cited by AI answer engines via Collimer’s own five-engine scan panel. The site’s name and URL are withheld: the value here is the mechanism, not calling out one small company. As of August 2026.

Measure where you stand.

Run a free scan