Docs As Code
FeaturesLong read

In-Context Learning From Docs vs Retrieval at Query Time

Timing and retrieval strategy shape cost, reliability, and freshness more than raw performance.

Senior Correspondent · · 10 min read
Cover illustration for “In-Context Learning From Docs vs Retrieval at Query Time”
Features · October 8, 2026 · 10 min read · 2,161 words

In-context learning from documents and retrieval at query time both ground a language model in outside knowledge without touching a single weight, and that shared goal is why people use the two terms as if they were synonyms. One method bakes a document into the prompt before the model ever runs; the other goes out and fetches a slice of a living corpus each time a question comes in, and those two timings shape everything downstream: cost, staleness, failure modes, and which systems are even worth building.

In-context learning from docs versus retrieval at query time

Document-grounded generation takes a document, places it directly into the input prompt, and runs the model once, with everything the model will ever see about that corpus already sitting in the context window at execution time. Koga and colleagues, writing in 2025, draw this line with unusual precision: document-grounded generation supplies a preprovided document embedded in the input at execution, while RAG retrieves dynamically from external sources before each generation. Both count as in-context learning: neither updates model parameters, and both work by changing what enters the query. But treating them as the same operation misses the part of the system that actually determines cost and behavior.

The mix-up isn't hypothetical. Hewitt and colleagues, replying in 2025, conceded the distinction, acknowledging that document-grounded generation and retrieval are handled differently in practice. They also reported something more unsettling: when they asked major LLM providers, including ChatGPT and Anthropic, how their systems actually process uploaded documents, the answers came back as generic, bot-generated responses. If the companies building these systems can't produce a clear answer about their own document-handling pipeline on request, it says something about how blurry the line still is in practice, even among people who should know it best.

Cost of each mechanism: tokens, time, and memory at scale

The mechanism gap turns into a cost gap fast. Document-grounded generation pays the full token price of the entire document on every single query, since the whole thing sits in the prompt each time. That difference compounds rather than stays fixed, because transformer attention scales quadratically with the length of the context: double the tokens and the computation needed roughly quadruples. A cache holding something on the order of a million tokens needs something in the neighborhood of 100 gigabytes of GPU memory per session, which turns long-context ICL into a hardware constraint as much as a line item on a bill.

Latency follows the same shape. The token cost of repeatedly feeding a full document into every prompt never amortizes.

Reliability degradation from excess context and retrieval

More context doesn't reliably produce better answers from either approach, and that matters because it cuts against the intuitive fix of "just give the model more to work with." Long-context ICL fails for architectural reasons. Positional encoding schemes like RoPE carry a long-term decay property: the mathematical similarity between tokens that are far apart in the sequence shrinks as the distance grows, which systematically lowers the attention weight the model assigns to information sitting in the middle of a long context. That's a property of how the attention mechanism computes relationships between tokens, not something a better-written prompt can route around.

Retrieval has its own failure mode, and it turns out to be worse on the generation side than on the retrieval side. Researchers have documented what's been called "Evidence Override," where the model successfully retrieves the correct evidence and then simply ignores it when generating its answer. In medical and multi-hop question-answering settings, this generation-side failure accounts for four to seven times more hallucinations than cases where retrieval itself failed to surface the right document. The system found the right needle and still answered as if it hadn't.

None of this amounts to a blanket case against long-context prompting, and the exception is worth taking seriously. CUT What does that suggest about where the real risk sits? Not in context length by itself, but in the complexity of what's being asked of that context.

Diagram: Evidence Override: Where RAG Actually Fails. Visualizes: Show the split between two failure modes in retrieval-augmented generation: retrieval failure (the system surfaces the wrong document) versus generation-side failure called 'Evidence…

How freshness and staleness land differently by mechanism

Beyond raw performance, the two mechanisms age in opposite ways. Retrieval pulls from a live source at the moment of the query, so as the underlying knowledge base drifts, a RAG system's answers drift gradually along with it, tracking the source. Document-grounded generation degrades in a different way. The document was frozen into the prompt the moment someone built it in, and from that point forward it answers from that frozen snapshot regardless of what's changed since, with no built-in signal telling anyone the information has gone out of date.

That asymmetry turns static documentation into a quiet liability for any AI-powered product built on top of it. Koga and colleagues note the same concern from the clinical side: contextual data that's overly complex or simply out of date introduces errors, and they stress that input data needs regular updating as a matter of practice, not an afterthought.

Documentation platforms built for AI-agent consumption treat this problem as the design center. Mintlify, for instance, is built around the idea that documentation functions as a living source of truth feeding directly into agent workflows, rather than a static file that gets dropped into a prompt once and forgotten. Static ICL context can't offer that property on its own. It goes stale the instant the underlying corpus changes, so platforms built around agent integration put heavy weight on keeping the documentation and the product continuously synced with each other.

Where each approach breaks down in high-stakes domains

Abstract tradeoffs turn concrete once real domains are involved, and the failure patterns in law and medicine look different enough from each other that the right mechanism genuinely depends on the setting. Stanford Law School's 2025 evaluation of leading RAG-powered legal AI tools found hallucination rates running from roughly one in five to one in three queries, including citations grounded in the wrong source and outright false statements of law. These aren't errors a casual reader would catch; spotting a misgrounded citation or a misstated legal rule takes the kind of domain expertise most users of a legal research tool are trying to buy their way out of needing.

Clinical pathology tells a more encouraging story for the opposite mechanism. Koga and colleagues show document-grounded generation performing well on constrained clinical classification tasks, where the governing document is something like the WHO CNS5 criteria, a small, curated, authoritative reference that doesn't get revised on a weekly basis. That stability is what makes document-grounded generation workable there: the corpus is small enough to fit comfortably in context, curated enough to trust, and stable enough that staleness isn't a live risk.

The two cases side by side yield a working principle. Document-grounded generation holds up when the corpus is small, stable, authoritative, and the task amounts to classification against a fixed schema. Retrieval becomes necessary once the corpus is large, heterogeneous, or updated frequently, though it demands a lot more engineering discipline to keep reliable, since retrieval quality, not just generation quality, becomes a second place for things to go wrong.

The open-source versus closed-source model divide

Part of why published results on this question seem to contradict each other traces back to a structural difference between model families. Closed-source models more often bring stronger long-context capabilities to the table, and they tend to perform better when given the full document directly, which makes document-grounded generation relatively more competitive on them than it would be elsewhere.

Becker and colleagues frame the underlying retrieval problem in a useful way: given a particular query and a particular model, the real question is which subset of a practically unbounded dataset actually matters for answering it. The answer to "which subset matters" shifts with the model's own capacity to make use of what it's given, so a result showing document-grounded generation beating retrieval on one model says little about what will happen on a model with a smaller context window and weaker long-range attention.

The practical consequence for anyone comparing benchmark numbers: a result can't be read without knowing which model family produced it. A benchmark showing long-context prompting winning on a strong closed-source model doesn't transfer automatically to a deployment running an open-source model with a tighter context ceiling. The model is as much a variable in this comparison as the mechanism is.

The ingest-time compilation paradigm that sits between the two approaches

Framing this as a binary choice between document-grounded generation and retrieval skips over a third design that recent research has started to take seriously: ingest-time semantic compilation, which separates the cost of indexing a corpus from the cost of answering a query in a way neither of the other two approaches does.

IndexRAG, from Bao and Shi, accepted to Findings of AACL-IJCNLP 2026, moves cross-document reasoning out of the query path and into the indexing path. Bridging facts, the connective reasoning needed to link information across multiple documents, get computed once, offline, during indexing, so that answering a query at inference time takes a single retrieval pass and a single call to the model.

A broader version of the same idea appears in work by Wild, Takahashi, and Uraki, titled "RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation" (arXiv 2608.20845, August 2026). They propose ingest-time semantic compilation, built on a two-layer substrate: incrementally maintained embeddings paired with atomic claims whose sourcing gets checked at compile time. They treat the result as a database object in its own right, with its own schema definition, its own rules for maintenance, its own migration path, and its own cost accounting, rather than as a cache that happens to sit in front of a model.

On a held-out sample of broadcast-interview transcripts, compiled claims won every cell in a 32-cell grid crossing budget against model: the compiled approach produced correct answers using a small token budget for retrieval, while the best competing chunk-based configuration needed a much larger budget to match it. Updating the compiled substrate incrementally costs far less than rebuilding it from scratch, and those updates track the source corpus with floating-point precision, so the substrate avoids the silent staleness that plagues static document-grounded generation. A separate controlled experiment (arXiv 2609.29661), testing against corpora that had been revised after initial indexing, found query-time reconstruction producing the correct value, source, and revision together in only a small share of trials, while the compiled substrate got it right consistently, and at a fraction of the read cost, with the gap widest on the hardest revision questions.

That said, the evidence deserves a measured read. The reviewer community treats ingest-time semantic compilation as strong preliminary evidence rather than a settled conclusion, partly because the headline accuracy numbers come from the same model family that performed the original extraction, and the 32-of-32 win rate hasn't yet been replicated on an independent pipeline built by a different team. The idea is promising.

When prompt caching changes the calculus for long-context ICL

The cost argument against document-grounded generation assumes the full document gets paid for on every query, and prompt caching breaks that assumption for a specific, well-defined set of corpora. When the underlying document doesn't change often, the cached context can be reused across many queries, so the expensive quadratic attention cost gets paid once. That changes the arithmetic meaningfully for a corpus that sits still.

For knowledge bases under roughly 200,000 tokens, full-context prompting with caching can beat the alternative of building out retrieval infrastructure from the ground up, since a vector database, an embedding pipeline, and the tuning needed to make retrieval reliable all carry real engineering cost of their own. The regime where caching wins has a clear shape: a small, stable, authoritative corpus, high reuse of that same cached context across many queries, and an update cadence the team controls. The regime where it breaks down looks just as clear: a large corpus, frequent updates, and context that churns faster than any cache can keep up with.

That boundary maps onto a distinction documentation teams already live with day to day. One sits comfortably inside the caching-friendly regime. The other pushes straight back toward retrieval, almost by definition, because the content itself refuses to hold still long enough for a cache to pay off.

Diagram: Which Mechanism Fits Which Corpus. Visualizes: Visualize the two-dimensional decision space that determines the right mechanism: corpus size (small vs.

Retrieval quality and the structure and cleanliness of what you index

Neither mechanism escapes a simpler, more basic constraint: the ceiling on performance is set by the corpus itself, not just by which mechanism sits on top of it. Most retrievers are built assuming a clean, well-organized corpus to search over. A retrieval system searching over redundant, poorly structured documents will surface noisy results no matter how well-tuned its ranking model is, and a document-grounded system fed a messy, contradictory source document will reproduce that confusion directly in its answers.

Getting the corpus itself right matters more than choosing between document-grounded generation and retrieval. The mechanism matters. What it's built on top of matters at least as much.

Sources

  1. Retrieval‐augmented generation versus document‐grounded generation: a key distinction in large language models
  2. Authors' reply: Re: Koga et al. Retrieval‐augmented generation versus document‐grounded generation: a key distinction in large language models
  3. Maximally-Informative Retrieval for State Space Model Generation
  4. Inference Scaling for Long-Context Retrieval Augmented Generation
  5. When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs

More in Features