One year ago this week, Anthropic launched Claude Sonnet 4.5 and tucked the most interesting detail in AI into the middle of the announcement. The launch had more going on than a benchmark bump. The model gathered some of its own training data using an agentic search engine that, per Anthropic, ran over 500 searches to synthesize a single training document.
Retrieval shaped the weights.
For a decade, retrieval was something you bolted onto a model at inference time. Enterprise RAG stacks exist because the model alone doesn’t know your product docs, your tickets, your contracts. But a model that spent training learning to search, judge sources, and cite arrives at runtime with very different habits than the passive reader most stacks were designed around.
The industry half-noticed. Coverage focused on the benchmark jump, 77.2% on SWE-bench Verified per Anthropic’s numbers, and on the coding revenue milestone. Weeks earlier, OpenAI had shipped GPT-5, which decides on its own when to search the web. No toggle required. Two frontier labs, two moves, same message: agentic search belongs to the model now.
A year on, the pattern has spread, and enterprise teams are feeling it in odd places. Models reformulate queries your retrieval API never saw. They cite your chunks with web-style formatting. Some skip retrieval entirely, answer from stale priors, and apologize when a user corrects them. The stack wasn’t built for any of this.
None of this kills RAG. It rewrites the stack’s job description, which is why this is the story worth your attention this week. Here’s what changed, why the weights matter, and five shifts your team can start on now.
The Agentic Search Engine Buried in the Sonnet 4.5 Launch
On September 29, 2025, Anthropic published its Claude Sonnet 4.5 announcement. The interesting part sat below the benchmark tables. Anthropic described an agentic search engine that could run over 500 searches to produce a single synthesized training document, citations attached. The engine gathered raw material from the web and turned it into instruction data. The model, in effect, helped build its own curriculum.
Two things make this more than a data pipeline trick.
First, the synthesized documents carried citations. Source attribution, the thing RAG teams spend whole sprints on, existed at training time. A model raised on cited documents has citation habits compiled into its behavior rather than prompted into it.
Second, the habit generalizes. A model trained by searching acts like a searcher at inference time. It formulates queries, judges weak results, reformulates, and decides when to stop. Anyone who has watched GPT-5 decide on its own to hit the web has seen the same instinct from the other direction. OpenAI removed the manual search toggle and let the model route itself.
Agentic search landed twice in the same year, once at training time and once at inference. The first changed what models are. The second changed what models do. Both change what your retrieval stack is for.
What Search-Native Training Changes
Start with the two flavors, because they get conflated.
Inference-time search is the familiar one. The model queries an external source mid-conversation, whether that’s GPT-5 hitting the web or an enterprise agent hitting your vector store. Training-time agentic search is the Sonnet 4.5 move: a retrieval system participated in constructing the training data itself, so search behavior got compiled into the weights.
The distinction matters for a practical reason. A model that merely has a search tool will use it when prompted well. A model shaped by search will expect retrieval to exist, will judge its quality, and will work around it when it’s bad.
That’s the new client sitting on the other end of your retrieval API.
It helps to remember why RAG won the enterprise in the first place. Menlo Ventures’ 2024 survey of enterprise generative AI found RAG behind 51% of production investments, ahead of fine-tuning at 31%. Teams didn’t pick RAG because it was elegant. They picked it because public models can’t see private data, and because baking knowledge into weights is brittle, expensive, and a security headache.
None of that has changed. A search-native model still can’t see your SharePoint, your Jira, or your contracts. Agentic search at training time fixed stale public knowledge. It did nothing for private knowledge. That gap is still yours to fill.
What changed is the interface. For years, retrieval teams tuned for a model that would consume whatever chunks it was given. The next generation walks in expecting a search engine, because that’s what raised it. Your stack’s job is no longer to hand the model knowledge. It’s to be a search engine the model already knows how to interrogate.
Same components, different contract.
5 Shifts Your RAG Stack Needs for Agentic Search
Practical version: what to do. Each shift comes with a first step you can finish this week.
1. Design for a Searcher, Not a Reader
Agentic search models don’t submit one query and wait. They loop: query, skim, reformulate, query again. In production, that means your retrieval endpoint needs to handle bursts of related queries, not lone lookups.
Three moves matter. Log full query sequences, not just individual searches, so you can see how the model explores. Budget latency per hop, because a model running five queries has less patience per query than one running a single lookup. And tune your rate limits so a burst of reformulations doesn’t lock out your best users.
First step: pull a week of query logs and count queries per answer. If the average is climbing past three, your stack is already serving a searcher, whether you designed for it or not.
2. Chunk Like a Web Page, Not a Database Row
Models raised on web search expect passages that stand on their own: a title, a source, a date, then the content. A floating 512-token paragraph with no provenance reads like a garbage result to a searcher. And searchers improvise when results are garbage.
Prepend context into the chunk text itself: document title, section path, last-updated date. Keep headings inside chunks instead of stripping them as boilerplate. A chunk that opens with “Refund Policy, Section 4.2, Vendor Handbook, updated March 2026” gives the model the orientation it would get from a web page, and it cleans up your citations for free.
First step: sample twenty chunks from your index and read them cold. If you can’t tell which document a chunk belongs to, neither can the model.
3. Put Permissions in the Query Path
Iterative search multiplies access events. If your stack filters permissions after retrieval, an agent running eight hops just gave itself eight chances to surface something it shouldn’t see.
Filter at query time, inside the retrieval call, using the end user’s identity rather than the agent’s. The OWASP Top 10 for LLM applications already treats vector and embedding weaknesses as their own risk class, and agentic search raises the stakes without changing the rule.
First step: run a test query as a low-privilege user and check whether restricted documents appear in intermediate retrieval steps, not just final answers. Intermediate leaks are where agent stacks get burned.
4. Evaluate Search Behavior, Not Just Answers
Most RAG evals still grade a final answer against a single gold query. That misses the failures search-native models produce: over-trusting bad chunks, ignoring good ones and answering from priors, or burning ten searches on a question that needed one.
Track four numbers instead: retrieval win rate (did the right chunk appear in the top k), abstention rate (how often the model admits the corpus doesn’t know), query depth per answer, and cost per question. Together they describe how the model uses your stack, not just whether the output looks right.
Vectara’s hallucination leaderboard has frontier models scoring in the low single digits on summarization tasks, which sounds solved until you notice those tests never touch your corpus. Behavior with retrieved context is a different measurement. RAGAS and citation-faithfulness checks get you closer.
First step: add query depth to your eval dashboard. It’s one counter, and it’s the cheapest signal you’ll ever add.
5. Treat Model Releases as Breaking Changes
The uncomfortable part: if search behavior lives in the weights, every model release potentially changes how your retrieval stack gets used. Not the API, the behavior.
Teams already lived this once with reasoning models, which changed the chunk sizes and query styles that worked overnight. Search-native models will do it again. Reformulation patterns, citation formats, and search-versus-answer decisions will drift between versions.
Pin model versions in production the way you pin library versions. Run your eval set against a new model before switching, with attention on query depth and abstention rate. Keep a frozen set of your weirdest production questions too, because release notes will never tell you that the new model suddenly prefers three narrow queries over one broad one.
First step: write down, on one page, which model versions your stack assumes. If the answer is “whatever the default is,” you’ve found the gap.
What Agentic Search Models Still Can’t Do
None of this makes the retrieval layer obsolete. In most enterprises it makes the layer more important, not less.
Agentic search models can’t see private data. Their training engines searched the public web. Your product docs, support tickets, contracts, and internal policies remain invisible to them, which is exactly why enterprise RAG took off in the first place.
They can’t ship into regulated environments unchanged, either. A model that autonomously searches the web is a non-starter for legal, healthcare, and finance teams whose data can’t leave the building. Those teams will keep running retrieval against local corpora, treating the model’s web instincts as a compatibility requirement rather than a feature.
And they can’t keep enterprise data fresh. Web search gets you the news. It won’t get you the pricing sheet that changed Tuesday or the ticket that closed this morning.
So the retrieval layer survives, but its identity changes. It stops being a patch over the model’s ignorance and becomes the model’s search engine: a real one, with permissions, freshness, and provenance. Teams that internalize this will build stacks the new models like using. Teams that don’t will spend their days wondering why the model keeps improvising around their chunks.
Where This Leaves Your Stack
One year ago, a search engine wrote part of Claude’s homework: 500-plus searches per document, citations included. This year, the model grades yours.
The summary is short. Agentic search moved inside the model, first at training time, then at inference time, and models now arrive with search instincts your stack has to serve. The five shifts follow from that: design for multi-query searchers, chunk with provenance, filter permissions in the query path, evaluate search behavior, and pin model versions like the breaking changes they are. None of it requires a rebuild. All of it requires treating the model as an active participant in retrieval rather than a passive consumer.
If your team is feeling this already, climbing query depth, strange reformulations, models skipping your chunks, we want to hear about it. Rag About It publishes weekly teardowns of exactly these shifts, from chunking standards to eval practices for agentic retrieval. Subscribe to the newsletter to get the next one in your inbox, and reply with the strangest retrieval behavior your stack has logged lately. The best stories become case studies, names optional, lessons public.
The model learned to search. Your stack gets to decide what it’s searching.



