Alok Upadhyay | May 2026 · Explainer
TL;DR: RAG is drawn as embed, retrieve, generate, and that drawing hides every decision that matters. Two properties decide whether the system is any good: whether the chunk containing the answer can be retrieved at all, and whether the answer that comes back is traceable to the chunks that were actually retrieved. The first is bounded at ingestion time, long before a query arrives. The second is a check you have to build, because the model will not volunteer it.
The framing that helps
A RAG system is not a model with a database attached. It is a search system with a generation stage bolted to the end, and it inherits every problem search has ever had: recall, freshness, ranking, permissions, deletion. The language model does not fix any of them. It does something worse, which is to paper over them fluently.
That framing has a practical consequence. Recall@k at the retrieval stage is a hard ceiling on the whole pipeline. If the passage containing the answer is not in the k chunks you hand the model, no amount of prompt engineering, reranking or model upgrade will recover it. Everything downstream can only lose information. So the money is spent in the places that raise that ceiling, and most of those are on the indexing path, where nobody is watching.
The indexing path
Chunk boundaries set the recall ceiling: nothing downstream can retrieve an answer the chunker split in half.
Chunking is the highest-leverage decision in the system
Chunking looks like preprocessing. It is really a decision about what the smallest retrievable unit of meaning in your corpus is, and it sets the recall ceiling directly.
Fixed-size chunking (512 tokens, 50 overlap, move on) is the default because it is trivial to implement and uniform, which the embedding model likes. It also splits tables in half, severs a heading from the paragraph it governs, and cuts a procedure between step 3 and step 4. When the answer is “steps 3 through 5”, and steps 3 and 4 live in different chunks, the retriever’s best case is returning half an answer.
Structure-aware chunking respects what the document already tells you: headings, list boundaries, table rows, function definitions, message turns. It is more work per source type, because a Confluence page, a PDF spec and a support thread have different structures, and the work does not transfer. It is usually worth it anyway, because the alternative is capping your recall on the exact queries users care about most, which are the specific ones.
Three details that recur:
- Overlap is a patch, not a fix. A 10-20% overlap catches answers that straddle a boundary by a sentence. It does nothing for an answer that spans two sections, and it inflates the index.
- Chunks need their context back. A chunk retrieved in isolation often reads as meaningless: “this is not supported in the enterprise tier” with no subject. Prefixing each chunk with its document title and heading path, at index time, costs almost nothing and makes both the embedding and the generated answer better.
- Small to retrieve, large to read. A useful pattern is to embed a small precise chunk but expand to its parent section when assembling context. Retrieval stays sharp; the model gets enough to actually answer.
Whatever you pick, keep the chunker versioned and deterministic. When retrieval quality moves, the first question is what changed, and “the PDF parser was upgraded” needs to be an answerable question.
The embedding model is a schema decision
Choosing an embedding model feels like choosing a library. It is closer to choosing a database schema, because every vector in your index is only comparable to vectors from the same model and the same version. Change the model and the existing index is not degraded, it is meaningless: you are comparing coordinates from two different spaces.
This means an embedding upgrade is a full reindex of the corpus, and for a large corpus that is a scheduled operation with a cost, not an afternoon. The sane procedure is to build the new index alongside the old one, serve from the old until the new one is complete and evaluated, then swap. What you must never do is write new chunks with the new model into the index the old model built. That produces an index where similarity scores are incomparable across a subset of rows, and the symptom is a retrieval quality regression that no single query reproduces.
Practical consequences worth deciding up front: stamp the model id and version on every chunk record so a mixed index is detectable rather than merely suspected; treat dimensionality as part of the storage plan, since it drives index size and memory; and be aware that any fine-tuning of the embedder has the same blast radius as swapping it.
Metadata is captured at index time, not invented at query time
Every chunk should carry, as indexed fields: source document id, version, owner, last-updated timestamp, source system, document type, and the access-control tags that govern it.
The access tags are the ones people get wrong, and they get them wrong by deferring. The tempting design fetches candidates first and checks permissions afterwards, or worse, retrieves broadly and asks the model to be careful. Both are wrong, for reasons I will come back to. Access tags have to be resolved when the chunk is written, so that the retrieval query itself is filtered, which means the ingestion pipeline needs a live view of the source system’s permission model. That is real work, and it is the work.
Freshness metadata earns its place too. A chunk that says “the current limit is 200” from a document last updated two years ago is worse than no chunk, because it is confidently specific. If the index knows the document’s age, the ranker can down-weight it and the answer can surface it.
Incremental update and deletion is where systems rot
The first build of an index is the easy part. Every RAG system that has been in production for a year has the same problem, which is that the index and the source of truth have drifted apart.
Full reindexing on a schedule is simple and, for a small corpus, correct. It does not scale, and it makes freshness a function of the batch window: a document edited an hour after the nightly run is wrong in your index for twenty-three hours.
Incremental ingestion driven by a change feed is the answer, and it has to handle three cases, not one:
- Insert is easy and is the only one anybody tests.
- Update is not an insert. Rechunking a modified document produces a different number of chunks with different boundaries, so the old chunks of that document must be deleted as a set, keyed on document id, before the new ones are written. Upserting chunk by chunk leaves orphans from the previous version, and those orphans are now stale content that looks perfectly valid to the retriever.
- Delete must propagate. A document removed at the source, or a document a user has lost access to, has to disappear from both indexes. Soft deletes that filter at query time are acceptable only if the filter is genuinely enforced everywhere, including in whichever debug path someone added last month.
The reason this rots is that failure is silent. Nothing alerts when a deletion event is dropped. The index just quietly keeps serving a document that no longer exists, and you find out when someone asks the assistant about a policy that was retired in March. An index-versus-source reconciliation job, running on a schedule and reporting counts and checksums, is cheap insurance and is almost never built until after the first incident.
Freshness against cost
Freshness is a knob with a bill attached, and it is reasonable to set it differently per source. A pricing page and an incident channel need minutes. An archived design document can tolerate a week. Tiering the ingestion cadence by source type is usually a better trade than either extreme: one global real-time pipeline is expensive, and one global nightly batch is wrong for the sources people ask about most urgently.
The query path
Two retrievers, because each one is reliably wrong about the other one’s easy cases. Retrieval sets the ceiling; generation can only lose information the retriever already found.
Query understanding, and the multi-turn problem
The user’s question is frequently not a good search query, and in conversation it is frequently not a question at all. “What about on iOS?” is meaningless as a retrieval query. It is only answerable in the context of the previous three turns.
So the first stage rewrites the incoming turn into a self-contained query. Small models do this well and cheaply. Alongside the rewrite, this stage can extract structured filters that belong in the query rather than the embedding: a date range, a product name, a document type. Filters expressed as metadata constraints are exact; filters expressed as hope that the embedding captures them are not.
Two cautions. Rewriting adds a hop to the latency budget and a new failure mode, since a bad rewrite loses information the original turn had. And generating multiple query variants raises recall at the cost of a wider fan-out downstream, which is a real budget decision, not a free win.
Hybrid retrieval, because dense alone loses specific queries
Dense retrieval embeds the query and finds nearest neighbours. It is very good at the thing lexical search is bad at: matching paraphrase, synonym and intent, so a query about “turning off notifications” finds a document about “muting alerts”.
It is unreliable in a specific and predictable set of cases, all of which happen constantly in enterprise and product search:
- Exact identifiers. Error code
E4021, ticketPROJ-3318, SKU numbers, config keys. Embeddings of rare tokens are weakly trained and near-identical strings sit close together in the space, so the retriever cheerfully returnsE4012. - Rare and novel terms. A product name coined after the embedder was trained has no meaningful representation.
- Negation and quantifiers. “Which regions do not support this” embeds close to “which regions support this”. Bi-encoders compress a passage into one vector, and the negation is one of the first things lost.
- Queries where the user already knows the wording. People who know the phrase in the document want that phrase matched, not approximated.
Lexical retrieval (BM25 and friends) handles all of those natively, because it matches terms. It fails equally predictably on vocabulary mismatch.
Run both, over the same chunk ids. The merge is usually reciprocal rank fusion, which combines by rank rather than by score and therefore does not require calibrating a cosine similarity against a BM25 score, two quantities with no common unit. A tunable weight between the two arms is worth having: identifier-heavy corpora want more lexical, prose-heavy corpora want more dense.
Reranking is where the accuracy is
The retrievers are bi-encoders and inverted indexes: fast, because the document side is precomputed, and approximate, because the query and the document never meet. A cross-encoder puts the query and a candidate chunk through the model together, with full attention between them. It is far more accurate about relevance, and far too slow to run over the corpus.
Which is exactly the funnel shape: retrieve a wide candidate set cheaply, rerank a short list expensively, keep a handful. The rerank stage is often the single largest quality improvement available for the least architectural change, because it is a drop-in stage between two things you already have.
Two things to hold on to. The reranker’s candidate list is capped by what fusion handed it, so reranking cannot repair a retrieval miss; it can only reorder a set that already contains the answer. And the rerank depth is a latency dial: cross-encoder cost is linear in candidates, so the choice of 50 versus 200 is a quality-latency trade you should make deliberately and measure, rather than inherit from a tutorial.
Context assembly under a budget
You now have a ranked list and a finite window. Packing it is not concatenation.
The token budget is shared with the system prompt, the conversation history and the space the answer needs. Chunks retrieved from near-duplicate documents waste it: the same paragraph from three versions of the same page is one piece of evidence, priced as three. Deduplicate on content, not just on chunk id.
Ordering matters more than people expect. The “lost in the middle” effect described by Liu et al. is that models attend more reliably to material at the start and end of a long context than to material buried in the middle. Sorting the reranked chunks so the strongest evidence sits at the edges is close to free and measurably helps. It also means a very long context is not a substitute for good ranking: stuffing 100 chunks in because the window allows it degrades the answer relative to 8 well-chosen ones, while costing more and taking longer.
Finally, each chunk goes in with a stable identifier and its source metadata attached, because the citation the model produces has to resolve to something. A citation the user cannot click is decoration.
Generation, citations and the check afterwards
Generation is constrained: answer from the supplied context, cite per claim, and say so when the context does not contain the answer. That last instruction is the one that needs reinforcing, because the default behaviour of a fluent model handed insufficient evidence is to produce a plausible answer anyway.
Then check the output. A separate, cheaper pass that asks whether each claim in the answer is supported by the chunk it cites catches the two failures that matter most: a claim with no support in the retrieved set, and a citation attached to a chunk that does not actually say the thing. Same evidence, second look. The generation step is optimising for a coherent answer; a targeted check is optimising for whether the answer is true given the evidence, and those are different objectives that should not be run in the same call.
What happens on a failed check is a product decision, so make it explicitly: revise, drop the unsupported sentence, or return the retrieved passages with an honest statement that the answer was not found. All three beat shipping a confident fabrication.
Cross-cutting concerns
Permissions are enforced at retrieval
This is the one that turns an embarrassing bug into an incident, so it deserves to be stated plainly.
The permission filter belongs in the retrieval query. Not as a post-filter on results, and never as an instruction to the model.
Post-filtering is subtly broken even when it appears to work: if you retrieve the top 20 and then remove the ones the user cannot see, a user with narrow access gets fewer results than a user with broad access, from the same query. Their result count, and the pattern of what is missing, is itself a leak. And the implementation only has to be skipped once, on one code path, to expose everything.
Asking the model not to use what you gave it is not access control at all. Once a passage is in the context window it has been disclosed: it can be summarised, paraphrased, alluded to, or extracted by a user who asks differently. Prompt instructions are a preference, not a boundary.
The correct shape is that the retrieval call carries the user’s resolved permission set as a filter on indexed access tags, so unauthorised chunks are never candidates. This is why access tags belong at index time. It also means permission changes need a propagation path: a user removed from a group must stop retrieving that group’s content, and if your tags are denormalised into the index, that is a reindex event for the affected documents.
One more: the logs. Query logs and retrieved-chunk logs contain content from documents, and they inherit the same access constraints as the documents. Debugging dashboards are a common back door.
Evaluation has two halves, and conflating them is the mistake
A single end-to-end answer quality score tells you that something is wrong. It does not tell you what, and the two possibilities need completely different fixes.
Retrieval quality is measured on the retrieval stage alone, against a set of queries with known relevant chunks. Recall@k is the one that matters most, because it is the ceiling: if the answer is not in the k retrieved, the rest of the pipeline is irrelevant. Precision and ranking metrics like nDCG tell you how well the reranker is ordering what it found. Build this set from real query logs rather than synthesised questions, because real queries are messier and uneven in ways that synthetic ones are not.
Answer quality is measured given the retrieved context. Faithfulness: is every claim entailed by the context? Citation correctness: does each cited chunk actually support the claim attached to it? Completeness: did the answer use the evidence that was there? Appropriate abstention: when the context genuinely lacked the answer, did the system say so?
The split is the whole point. A bad answer from good retrieval is a generation or prompting bug. A bad answer from bad retrieval is a chunking, embedding or ranking bug, and no prompt change will touch it. Teams that track only the end-to-end number spend weeks tuning prompts against what is actually a chunk boundary problem.
Two supporting habits. Keep a small, curated regression set that includes the specific queries that have broken before, and run it on every index rebuild and every model change. And log the retrieved chunk ids alongside every answer, permanently, because without them a complaint about a wrong answer cannot be diagnosed at all.
The latency budget
RAG puts more sequential stages on the request path than a search box does, and each one is someone’s opportunity to add 40 ms. Assign a budget per stage and enforce it: query rewrite, retrieval (both arms in parallel, so the budget is the slower one), fusion, rerank, assembly, generation, check. The ones that tend to blow their budget are the rerank depth, which is linear in candidates, and the grounding check, which is an extra model call that arrives after the user is already waiting.
Streaming the answer hides generation latency but not retrieval latency, because nothing can start until the context is assembled. So the pre-generation stages are the ones with the real constraint, and any caching you do (of rewrites, of retrieval results for repeated queries) has to be keyed on the user’s permission set, or you have rebuilt the leak you just closed.
Failure modes
| Failure | What you see | Where it actually is |
|---|---|---|
| Retrieval miss answered anyway | Confident, fluent, wrong | Generation prompt has no abstention path; recall@k untracked |
| Chunk boundary splits the answer | Half-correct answers on procedural or tabular queries | Chunking strategy, set at ingestion |
| Embedding drift after a model swap | Broad quality drop no single query reproduces | Mixed model versions in one index |
| Stale or deleted document served | Answer cites something retired months ago | Change feed dropping deletes; no reconciliation job |
| Permission leak via retrieved context | User sees content they should not | Filtering after retrieval, or asking the model to withhold |
| Lost in the middle | Evidence was retrieved and ignored | Context ordering and over-stuffing |
| Citation does not support the claim | Looks sourced, is not | No grounding check; citation correctness unmeasured |
The first and the fifth are the two to design against from the start. The first is the one users notice and stop trusting you for; the fifth is the one that ends up in a post-incident review.
What to instrument
recall@k on a fixed query set, per index build;
this is the ceiling on everything else
abstention rate how often the system declines, and whether
those cases really lacked evidence
citation accuracy share of citations that support the claim
attached to them, sampled and checked
index lag source updated-at vs. indexed-at, p50 and p99,
reported per source system
orphan count chunks whose parent document no longer exists;
should be zero and rarely is
stage latency p99 per stage against its budget, not end to end
empty results queries returning nothing, or nothing above
the score floor; the leading indicator of a
coverage gap
permission denials retrievals filtered by ACL, by user segment;
a sudden drop means a filter stopped firing
The last one deserves the emphasis. A permission filter that silently stops applying produces no errors, no latency change and no user complaints. What it produces is more results, which looks like an improvement. Alerting on a drop in filtered retrievals is one of the few ways to catch it before someone else does.
The one-sentence version
Build the search system properly (chunk on structure, index the permissions, keep the deletions honest, retrieve both lexically and densely, then rerank), and the generation stage becomes a formatting problem; skip that work and the language model’s only contribution is making the failures harder to see.