The most interesting story in AI right now isn’t a model release. It’s been sitting in the engineering docs of the fastest-growing AI products for a year, and most teams building enterprise RAG haven’t read it. The agents that write our code, the ones churning through pull requests at 2 a.m., mostly don’t use the retrieval pipelines we spend months perfecting. They grep. And what they do instead is the best evidence anyone has on RAG for code retrieval.
Anthropic’s Claude Code documentation says the team tried embeddings-based retrieval early on and moved to agentic search because it performed better on real repositories. Cline, one of the most-installed open-source coding agents, ran the same experiment in public: it shipped vector search over code chunks, watched most users ignore it, and rebuilt the feature around a tree-sitter code outline plus the model’s own search loop. Cursor and GitHub Copilot kept embeddings, but only as one path in a hybrid, with the agent doing its own searching on top.
That should stop you cold if you build retrieval for a living. Coding agents are the largest deployment of retrieval-augmented AI in production, and the argument over RAG for code retrieval is the most evidence-rich retrieval debate the field has ever had. These agents answer questions grounded in a private corpus that changes daily, which is exactly the problem enterprise RAG exists to solve. The best-resourced teams in the industry, with more usage data than anyone, keep landing on the same design: hand the model a map, give it cheap search tools, and let it iterate.
Before you torch your vector database, the story is more interesting than another round of “RAG is dead.” Code is an extreme case of retrieval, with exact-match signals, hard structure, and instant verification. Your support docs are not code. But the patterns coding agents settled on translate almost one to one, and most enterprise stacks skip them entirely. Here’s why embeddings lose at RAG for code retrieval, how agentic search works under the hood, five patterns worth copying, and where embeddings still earn their keep.
Coding agents are the biggest retrieval systems in production
The numbers around coding agents have gotten silly. The Information reported in mid-2025 that Cursor passed $500 million in annualized revenue, roughly two years after launch. Anthropic has called Claude Code the fastest-growing product in its history. OpenAI shipped a dedicated coding model, GPT-5-Codex, and built its cloud agent business around it. GitHub turned Copilot into an agent that takes an issue and opens a pull request. When Claude Sonnet 4.5 launched in September 2025, it resolved 77.2% of issues on SWE-bench Verified, the benchmark where an agent gets a repository and a failing test.
Every one of those products lives or dies on retrieval. The model wasn’t trained on your repository. To fix a bug, it has to find the right file, the right function, and the right caller in a corpus it has never seen, one that changed since yesterday. That’s a RAG problem wearing a hoodie. And the feedback loop is brutal. Retrieve badly, and the agent invents an API that doesn’t exist, and the tests fail loudly, in front of everyone. Coding agents are the most honest test bed for retrieval design we have.
Anthropic’s Economic Index, a running study of millions of Claude conversations, keeps showing coding and technical work as the largest category of model usage. So when the engineers behind these agents publish what they’ve learned about RAG for code retrieval, enterprise RAG teams should read it like opponent game film. The conclusions line up more than you’d expect.
Why chunk-and-embed breaks RAG for code retrieval
Standard RAG chunks a document, embeds the chunks, and returns the k nearest neighbors of the question. For prose, that holds up. RAG for code retrieval breaks every assumption in the pipeline.
Chunking cuts functions mid-body, so the chunk that mentions retryBackoff usually doesn’t contain its definition. The nearest neighbors of a code chunk are other chunks that look similar, and looking similar is not what anyone needs. You need the interface it implements, the caller three files up, the config key it reads. Vector similarity is a poor proxy for dependency.
Code is also full of exact-match gold: symbol names, error codes, import paths. If the ticket says “retryBackoff is wrong after a 429,” grep finds that string in milliseconds, in every file, forever. An embedding search might return three functions that feel like retry logic. And this isn’t unique to code. On the BEIR benchmark, BM25, a keyword scorer older than most machine learning engineers, beat most dense retrievers on out-of-domain datasets (Thakur et al., 2021). Code questions are often identity questions, and identity questions want exact match.
Production teams found the same thing. Anthropic’s docs describe moving away from embeddings because agentic search with ripgrep performed better on real repos. Cline published an engineering post titled “We don’t use RAG” explaining that its @codebase feature, an embeddings index, saw weak engagement because Claude’s own search loop, running search_files and read_file over and over, simply did better. Cline rebuilt the feature as a tree-sitter outline, a parsed map of every file’s functions and symbols, cheap to build, with no vector math at all. The map orients the agent. The agent finds the rest itself. If you build RAG for code retrieval, that post is worth ten minutes of your week.
How agentic search actually works
So what does RAG for code retrieval look like when the model drives the search? Strip the branding and agentic search is a loop. The agent reads a task, guesses where the answer lives, runs a search, reads the results, and either answers or refines. The tools are almost embarrassing in their simplicity: list files, grep a regex, read a file, search git history. No ingestion pipeline required.
The model itself is the reranker. A single-shot vector search commits to one interpretation of the question before seeing any evidence. The agent updates its interpretation with every tool call. That iteration is the whole trick.
Anthropic’s engineering team laid out the philosophy in a September 2025 post on context engineering. Just-in-time retrieval beats just-in-case ingestion, so let the agent pull facts when it needs them rather than pre-loading everything. The same post pushed compaction, sub-agents, and writing notes to files as core habits for long sessions.
None of this is free. One agentic search can burn dozens of tool calls and several minutes to answer a single question. Coding agents get away with it because a merged pull request is worth real money. A customer-facing chatbot with a two-second SLA cannot run forty search rounds per query. So you gate it. More on that below.
Five patterns worth copying
You don’t need to rebuild your stack around ripgrep to learn from any of this. The patterns that won RAG for code retrieval translate to enterprise RAG. Here are the five.
1. Give the model a map before the search
Before Claude Code searches anything, it reads CLAUDE.md, the project’s orientation file: build commands, architecture notes, conventions. Cline’s tree-sitter outline does the same job, a structural digest of the codebase without the file bodies. A cheap map beats clever chunking. The map tells the model where to look.
The enterprise equivalents are sitting in your docs already:
- A generated overview of your documentation tree, one paragraph per section
- An entity page listing your products, their APIs, and their common error codes
- A changelog digest so the model knows what changed this quarter
Build the map once, update it on a schedule, and spend the first few hundred tokens of every session orienting the model instead of letting it guess.
2. Put exact match in the retrieval path
Most enterprise stacks treat keyword search as legacy. That’s backwards. Error codes, product names, policy numbers, ticket IDs: these are grep questions in disguise, the same identity questions that dominate RAG for code retrieval. Vector search answers them badly, returning semantically similar text instead of the string the user quoted.
Run hybrid retrieval, fuse the results, and rerank. BM25 plus dense vectors plus a reranker is the boring default for a reason. If your corpus is mostly prose, exact match will quietly rescue a fifth of your queries, and your semantic benchmarks won’t show it unless you look for it.
3. Make retrieval a loop, with a budget
Single-shot top-k retrieval handles most queries fine and fails the multi-hop ones completely. Coding agents handle the hard cases by iterating. Your RAG system can too, on a leash.
Gate the agentic path with something cheap: query length, a follow-up flag, a one-line classifier. Cap the loop at five rounds, not forty. When the budget runs out, return a partial answer with citations rather than nothing. The deep research products everyone envies are doing exactly this, just with bigger budgets. Your support bot needs the pattern. It just doesn’t need the forty rounds.
4. Judge retrieval by outcomes
When a coding agent retrieves the wrong file, the tests fail. The compiler doesn’t care about recall@5. That’s why coding agent retrieval improved so fast: verification was automatic and unforgiving.
Enterprise RAG has the same signals, they’re just slower. Ticket resolution. Follow-up question rate. Whether a human opened the cited source. Whether the drafted answer got edited before sending. Instrument those. Token overlap with a gold answer was never the goal, it was a proxy, and the coding agents showed how much better the real thing works.
5. Keep memory outside the context window
Claude Code has CLAUDE.md. Cline has a memory bank. Anthropic’s context engineering post recommends agents write notes to files as they work and reread them later, rather than trusting the context window to hold everything. Compaction summarizes old turns so long sessions don’t degrade.
The translation for RAG products is direct. Persist what the session learned: user preferences, resolved entities, documents already read and rejected. Long research sessions should write their own running notes, or they’ll repeat searches and contradict themselves by turn twenty. The corpus is half of retrieval. The session is the other half.
Where embeddings still win
None of this means vector search is finished. When a user asks “how do I handle customers who won’t pay,” there is no string to grep. Paraphrase questions over prose are what embeddings were built for. Enterprise corpora are mostly prose. Multilingual search, semantic dedup, finding the paragraph that means the thing without saying the thing: those stay dense.
Notice that even the coding tools refused to go pure keyword. VS Code’s Copilot codebase indexing combines embeddings with a reranking pass. Cursor kept its vector index for @codebase and layered agentic search on top for the agent. The pattern that keeps winning is hybrid: keyword for identity, vectors for meaning, a reranker to sort it out, and an agent loop for the questions that refuse to die in one round.
The divide is the corpus, not the hype. RAG for code retrieval runs on exact match and structure. RAG for prose runs on meaning. So don’t delete the vector index, but stop treating similarity search as the entire system. Your corpus isn’t code, but it has more structure than your chunker admits: headings, tables, ticket metadata, timestamps, authors. The coding agents looked hard at their corpus, found its structure, and built retrieval around it. Most enterprise stacks never looked.
The takeaway
The past year of RAG for code retrieval has been one long public experiment. The agents churning through pull requests at 2 a.m. ripped out single-shot vector search and replaced it with maps, hybrid paths, loops, outcome checks, and external memory. They published the receipts in their docs and engineering posts.
You don’t need their budgets to copy the pattern. Map your corpus and put it in the system prompt. Add an exact-match path next to your vectors. Gate a five-step agentic loop for the hard queries. Instrument outcomes instead of proxies. Write session notes to disk. The map and the keyword path alone are an afternoon of work, and they’re the cheapest accuracy gains on this list.
We track shifts like this every week at Rag About It, because retrieval design is moving faster than the stacks built on top of it. Subscribe to the newsletter to get the next one in your inbox, and if your team is building or documenting a RAG system of your own, come tell us what you’re shipping. That’s the work we do.



