Watch a chatbot answer a hard question and you can spot the moment it starts inventing. The first sentence is clean. The second holds up. Then comes the third, where the model quietly drops its evidence and improvises a policy detail or a citation that doesn’t exist. Users rarely catch the switch.
The model never announces it. But it knows. Kadavath et al. showed in 2022 that uncertainty leaves traces in a model’s hidden states while it generates. A guess looks different from a recall under the hood. Most production RAG stacks never read those traces. KAIST’s CONDA framework does.
The standard answer is to retrieve for everything. Every query gets embedded, matched against a vector index, reranked, and stuffed into the prompt, whether the model needed help or not. That burns latency and budget on easy questions. Worse, it misses the failure mode that matters most: mid-answer drift, where the model wanders off its context and states a wrong fact at full confidence.
A team at KAIST, the Korea Advanced Institute of Science and Technology, published the framework this week. The CONDA framework flips the default: instead of retrieving by habit, it reads the model’s confidence in real time, token by token, and fires a search or retrieval call only when that confidence drops. The generation loop becomes its own router. No intent classifiers, no keyword rules. Just the model’s internal signal telling the pipeline when it needs backup.
The team’s numbers deserve a close look. Per the paper, CONDA caught roughly 9 in 10 hallucinated spans mid-generation, added single-digit milliseconds of overhead per token, and cut retrieval calls by more than half on workloads where most questions were easy. Fewer wrong answers and a cheaper pipeline at the same time. That trade is rare. This post covers the mechanics, the real cost of the retrieve-everything habit, and five takeaways you can test on your own stack.
What KAIST’s CONDA Framework Actually Does
Confidence you can read mid-sentence
Most uncertainty measurement is expensive. Semantic entropy, the approach Farquhar and colleagues at Oxford published in Nature, generates several answers to the same question and measures how much they disagree. It works. It also multiplies inference cost by the number of samples, and it still only runs after generation is done.
The CONDA framework takes the cheap path: it watches a single forward pass. As each token comes out, a lightweight probe reads the model’s internal states and produces a calibrated confidence score for that token and the span it belongs to. The probe trains on a few hundred labeled examples of correct and incorrect generations. That’s a weekend of annotation work, not a fine-tuning project. The base model doesn’t get retrained at all.
In a RAG setup, the interesting part is what happens when the score dips. Generation pauses. The pipeline builds a retrieval query from the original question plus the partial sentence, fetches evidence, and the model continues writing with the new context in view. Trigger and fix live in the same loop. The signal works just as well for agents that call web search: fetch when confidence wobbles, not on every turn.
What the numbers show
The KAIST team evaluated the CONDA framework on open-weight models across standard question-answering benchmarks. Results reported in the paper:
- Around 90% of hallucinated spans flagged before generation finished, with false positives low enough that easy answers stayed uninterrupted.
- Single-digit milliseconds of overhead per token on models in the 7B range, too small to matter for chat latency.
- Retrieval calls cut by more than half on workloads where most questions didn’t need external evidence.
- Final answer accuracy up several points over an always-retrieve baseline, partly because less stuffed context meant fewer distractions.
One honest caveat: these are the team’s benchmarks on their chosen datasets. The pattern held across every model they tested, but reproduce it on your own traffic before trusting a threshold.
Why Retrieve-Everything RAG Runs Hot and Still Misses
Run the cost math for a small deployment. Say your support bot handles 200 queries a day. Each one triggers an embedding call, a vector search, a reranker pass over 50 candidates, and a completion call with a prompt bloated by context. Most of those questions are things the model could answer from memory, or from three chunks instead of thirty. You pay for the full pipeline anyway, every single time.
Then there’s the distraction problem. Liu et al.’s “Lost in the Middle” study found that models attend most to the start and end of a context window, and lose accuracy when relevant information sits buried in long, padded context. Stuffing chunks the model never needed actively hurts answer quality. Retrieval isn’t free even when it’s correct.
But the real gap is timing. In a standard stack, retrieval happens once, before generation starts. After that first token, nothing watches the model. It can drift off its context at token 200, and nothing notices until a validator or an angry customer catches it downstream. A post-hoc check tells you an answer was wrong after someone already read it.
The CONDA framework closes the timing gap by changing the default. Retrieve on doubt, not on schedule. The KAIST results suggest most traffic never doubts, which is where the savings come from, and the traffic that does doubt gets caught in the act.
5 Takeaways for Your RAG Stack
1. Treat confidence as a routing signal
Most stacks route with heuristics: keyword match, an intent classifier, a follow-up flag. Those misfire in both directions and need constant gardening. Token confidence updates with every generation step and needs nothing between model versions except recalibration.
You don’t need to adopt the CONDA framework wholesale to test the idea. Logprobs from a standard API give a coarse confidence read. Log per-token confidence on a sample of real traffic for a week, then check what your low-confidence spans line up with: support tickets, thumbs-down ratings, whatever feedback you already collect. If the lows match the complaints, you’ve found a router. If they don’t, you saved yourself a rebuild.
2. Catch the guess at the token, not the answer
Validators that run after generation are popular because they bolt on easily. They also do nothing for the reader. The answer was already wrong when it hit the screen.
Mid-generation detection moves the failure point. Instead of “we caught 90% of hallucinations in QA review,” the story becomes “the user never saw the guess.” If mid-generation intervention is out of reach for your stack this quarter, run the confidence check per sentence and regenerate flagged sentences before display. It’s a weaker version of the same principle, and it ships in a sprint.
3. Spend the savings where they matter
If 40 to 60 percent of your queries stop hitting the index, that budget doesn’t vanish. Move it. The hard 15 percent of queries, the multi-hop ones that need three documents compared, are where retrieval quality differentiates you. A team that stops paying for single-hop lookups on easy questions can afford two-hop chains and a real reranker on the questions that earn them.
4. Calibrate per corpus, then recalibrate
Confidence scores don’t transfer. A probe calibrated on Wikipedia-style trivia will be overconfident on your legal contracts and underconfident on your support macros. Calibrate on a few hundred labeled examples from your own domain, not a public dataset.
Then keep doing it. If your corpus changes weekly, and most do, recheck the threshold monthly at minimum. A threshold tuned in March is quietly wrong by June.
5. Design the honest fallback
Sometimes confidence stays low even after retrieval returns good context. That should end in “I don’t know, here’s who does,” not a second confident guess. Abstention feels like failure to product teams. To users it reads as trust. A bot that admits it can’t verify a refund policy and hands off to a human keeps more customers than one that invents a policy number politely.
The CONDA framework treats abstention as a valid outcome rather than an error case, and your design should too. Build the handoff path before you build the router.
The Catches Before You Rewire Anything
Access is the first catch. The CONDA framework reads hidden states, which means open-weight models: Llama, Mistral, Qwen, and friends. On closed APIs you’re limited to logprobs, a weaker signal that still beats nothing but won’t match the paper’s numbers.
Calibration data is the second. You need labeled examples of correct and incorrect generations from your domain. A few hundred is enough. They still have to be yours, because public datasets calibrate the probe for someone else’s content.
The third catch gets less press: mid-generation retrieval stalls the stream. The user watches a sentence stop mid-word while the pipeline fetches evidence, and two seconds of silence feels long in a chat interface. Plan for it. Buffer the stream, show a “checking sources” state, or intervene at sentence boundaries instead of token boundaries.
And confidence can’t fix bad retrieval. If the index returns outdated or wrong documents, the model generates a confident, well-supported wrong answer that no probe will flag. Garbage in, confident garbage out. Keep your data freshness work going no matter how good the router gets.
Where This Leaves Your Stack
Back to that third sentence. In a confidence-guided loop, it stalls mid-write because the system is working, not failing. Retrieval fires, evidence lands, and the answer continues from the source instead of from memory. The user sees a brief pause and a citation instead of a fluent mistake.
The CONDA framework makes one simple move: retrieval becomes a decision the model makes with every token, not a reflex the pipeline runs for every query. KAIST’s numbers say you get fewer hallucinations and lower operating cost at once. For a small team, that’s the whole game. Better answers without a bigger bill.
If you run one experiment from this post, make it the confidence log. Sample real traffic, record per-token scores, and see what your lows actually predict. It costs an afternoon, and it tells you whether a confidence-guided router belongs on your roadmap.
We break down RAG research and production patterns every week at Rag About It. Subscribe to the newsletter if this one was useful, and send us your confidence-log results. The interesting ones become future posts.



