A few weeks ago, an engineer at a ten-person AI startup walked me through something odd. Their support copilot had answered customer questions at better than 90% accuracy for two quarters. Then, over about six weeks, the answers started drifting. The model hadn’t changed. Neither had the prompts or the chunking. Retrieval scores still looked healthy.
The problem wasn’t in the stack. It was in the corpus. RAG data quality, not model quality, had slipped. When the team traced its low-rated answers back to their sources, a pattern showed up: the web pages ingested months earlier had quietly turned over. Community threads that once held the best practical answers were filling with machine-generated replies. SEO farms had moved in. The retriever kept serving those pages because they were long, keyword-dense, and exactly on topic. Relevance was high. Trustworthiness was not.
That’s the small-scale version of the biggest story in AI right now. Researchers at the University of Waterloo and Colby College audited roughly 31 million web pages and estimated that about one in ten carried clear signs of LLM generation, with another large slice merely suspected. A Nature paper by Ilia Shumailov and colleagues showed what happens when models train heavily on model output: distributions flatten, the tails collapse, and each generation gets a little worse. Meanwhile, the clean data is getting fenced off. Reddit reportedly takes about $60 million a year from Google for access to its corpus, and the Financial Times reportedly signed a licensing deal with OpenAI worth more than $250 million.
If you build RAG systems, this is your problem too, not something the labs will solve for you. Retrieval augmented generation is only as good as what the retriever finds, and the retriever has no idea a page was written by a machine for machines. Retrieval quality used to mean clean chunking and decent embeddings. Now RAG data quality is the whole game: what’s in the corpus, where it came from, and whether anyone should believe it. The five defenses in this post turn your retrieval layer into a trust layer: source tiering, provenance metadata, ingestion screening, junk-aware evals, and a bias toward data you own. None of this needs a big budget. Most of it can ship in a week.
The Junk Flood Meets the Data Gold Rush
Three forces are colliding, and each one lands on your index.
Generation got nearly free. A content farm can publish a thousand keyword-targeted articles a day for the cost of API calls. The people running these operations test their output against search engines, and increasingly against answer engines. Reddit moderators have been publicly drowning in LLM comments for more than a year. The Waterloo and Colby College audit is just the cleanest measurement of a trend everyone in the industry has seen firsthand.
The incentives point the wrong way. Answer engines send traffic to whatever they cite, which has spawned a whole scene of SEO aimed at machines instead of readers. Write for the retriever, get retrieved. That’s the business model now.
And the good data is getting fenced. Cloudflare turned on Pay Per Crawl in July 2025 and started blocking AI crawlers by default, letting publishers charge per request. Major publishers sign licensing deals instead of getting scraped. Shayne Longpre’s team at MIT documented the squeeze from the other side in 2024: crawler restrictions rising across the biggest open corpora while machine-generated pages grew fastest where crawlers could still go.
RAG systems sit right on that fault line. Most enterprise stacks mix owned documents with community sources and open web pages. A Stanford RegLab study led by Joseph Dahl put hard numbers on why grounding matters in the first place: general-purpose models hallucinated on 58% to 82% of legal queries without it. Grounding was the fix. But grounding on polluted sources just gives you confident nonsense with citations attached.
Why Your Retriever Falls for Synthetic Pages
Embeddings measure semantic proximity. They don’t measure credibility. Those are different axes, and machine-generated junk is optimized for exactly the one your system scores on.
Synthetic content is, by construction, a near-perfect match for your queries. It copies the question’s phrasing. It’s structured with clean headings and lists, which parsers turn into tidy chunks. It carries fresh dates, so freshness boosts push it up the ranking. And it replicates: one template produces hundreds of paraphrased variants, enough to crowd a top-k list so completely that the single good source never surfaces.
Dedup doesn’t save you. Exact-match and shingle-based dedup catches copies, not paraphrases. Cross-encoders don’t save you either, because a cross-encoder rewards the same topical fit the embedding rewards.
So the failure stays quiet. Recall@k holds. Faithfulness holds, in the narrow sense that the answer follows from the retrieved context. The system is internally consistent and externally wrong, and nothing in your dashboard blinks. That’s why RAG data quality gets fixed at the data layer, not the model layer. You cannot prompt your way out of a poisoned corpus.
5 Defenses That Make RAG Data Quality Real
You earn this at ingestion and at query time. Here’s what it looks like in practice.
1. Tier your sources and write the policy down
Sort every source into three tiers. Tier 1 is what you own or license: product docs, the knowledge base, support tickets, official APIs. Tier 2 is vetted external sources, say a subreddit with active moderation or a vendor forum where real employees answer. Tier 3 is everything else.
Then make the policy code. Store the tier as metadata on every chunk, filter retrieval by tier where the query allows it, boost tier 1, and cap tier 3 at one or two slots in the final context. If a tier 3 page does get used, pass the tier into the prompt so the model can hedge in the answer.
Writing it down matters because it makes the policy reviewable. When someone proposes ingesting a new crawl, the tier policy forces the right conversation: who is this source, and why do we believe it? That conversation is where RAG data quality starts.
2. Stamp provenance on every chunk
Every chunk should carry the fields you’d want during an incident: canonical URL, first-seen timestamp, publish or last-modified date, author or organization, license, and tier. At the chunk level, not the document level, because one page can mix quoted authority with its own invention.
Provenance pays off in three places. Freshness decay becomes possible, with different half-lives for different source types. Incident triage gets fast: when a customer complains, you can query which sources taught the system that answer. And citations start to mean something, because your answer UI can show tier and date instead of a bare link.
The cost is lower than people assume. Parsers like Docling, Unstructured, and Apache Tika already extract most of these fields, and vector stores like Qdrant, Weaviate, and pgvector hold them as payload. Call it an afternoon of work per source type, not a platform project. Provenance is what turns RAG data quality from a guess into something you can query.
3. Screen for machine-generated content at ingestion
Don’t lean on GPT detectors. They’re unreliable on short text and biased against non-native writers. Use them as one signal among several, and score at the domain level instead. How many articles does this site publish per day? How old is the domain? What’s the ratio of outbound citations to word count? Does every author bio lead to a 404? Do the same paragraph templates recur across posts?
Score new domains before any full crawl. Anything suspicious goes into a quarantine bucket that only a human can promote. You don’t need perfection. You need the obvious farms to never reach the index.
One honest caveat: watermarks won’t save you. Google’s SynthID covers Google’s models, and most of the junk comes from unwatermarked open weights. Detection has to be statistical and behavioral, which is why those domain-level features matter more than text-level classifiers. Screening is the part of RAG data quality with no shortcut.
4. Put junk in your evals
Your eval set probably reflects the index you wish you had. Pollute it on purpose. Write 20 to 50 machine-generated decoy documents aimed at your golden queries, deliberately plausible and confidently wrong. Inject them into a staging index and run the normal suite. The metric is simple: what share of answers cite a decoy? Target zero, and treat any nonzero number as a release blocker.
Then watch production. Tag every answer with the tiers of its cited sources and chart the share of tier 3 grounding over time. If it climbs month over month, your corpus is rotting, and you’ll see it weeks before support tickets tell you. RAG data quality is a number you watch, not a box you check once.
The standard tools get you most of the way there. RAGAS-style faithfulness and answer relevance still matter. LangSmith or Langfuse traces will carry tier metadata if you put it there. The new metric, source quality, is one nobody’s framework ships for you, so you build it.
5. Lean on data you own
For a small team, open web pages are the worst ROI in the corpus. They’re noisy, they rot, and the good ones are being fenced off. Weight what you control instead: your docs, your anonymized support tickets, your release notes, licensed feeds, partner APIs. Owned data is where RAG data quality comes cheapest, because nobody can pollute it but you.
There’s a second move people miss. Make your own content machine-legible, both for your RAG stack and for the agents your customers run. Clean semantic HTML, an llms.txt file per Jeremy Keith’s proposal, real publish dates, per-section anchors. The same hygiene that makes your docs retrievable for your system makes them citable by everyone else’s.
The strategic read is simple. Models are turning into commodities you rent, and your competitor can rent the same one. A curated, provenance-clean corpus is not rentable. When a prospect asks what makes your RAG product defensible, “our data layer” is a better answer than “we picked a good model.”
Where to Start This Week
You don’t need a platform team for any of this. A first pass at RAG data quality fits into one week:
- Days 1 and 2: Trace your 50 most-used queries back to the sources your retriever is really pulling. Label each source tier 1, 2, or 3. Most audits find the problem in the first afternoon.
- Day 3: Add tier and provenance fields to ingestion for your top three source types.
- Day 4: Add the tier filter and boost to retrieval, plus a tier cap in the context builder.
- Day 5: Write 20 decoy documents, get a baseline for how often the system cites them, and put that number on the dashboard next to latency.
The economics favor acting now. Wrong answers cost you support escalations and churn, and they’re expensive to debug after the fact. A trust layer costs a few engineer-weeks and turns a silent failure mode into a chart you can watch. For a company under 10 people, that trade is easy.
The Takeaway
That engineer’s copilot is back above its old accuracy numbers. The fix wasn’t a new model or a bigger context window. The team tiered its sources, stamped provenance on every chunk, quarantined two suspect domains, and added a decoy set to the eval suite. Total effort: about two weeks, from people who also have a product to ship. It was RAG data quality work, start to finish.
The web will keep getting noisier and clean data will keep getting pricier. That direction won’t reverse on its own. Teams that treat the corpus as an asset, with policies, metadata, and tests, will own systems people trust. Teams that treat ingestion as plumbing will keep debugging confident nonsense at 2 a.m.
If you’re building production RAG, we cover these data-layer problems every week at Rag About It. Subscribe to the newsletter for implementation breakdowns like this one, and check out our guides on pipeline design and evaluation if you’re mid-build. Your retriever answers to whoever feeds it. Feed it well.



