The starter system from RAG from Scratch works — on the demo questions. Ship it, and you’ll meet its failure mode: it fails silently. The answer reads fluently while the retrieval quietly missed, and the model filled the gap with something plausible. Nobody notices until a customer does.
The production upgrade isn’t one trick. It’s a loop: better chunking, better retrieval, reranking, and — the part most teams skip — an eval harness that tells you which stage to fix.
1. Chunking: the highest-leverage boring work
Retrieval quality is bounded by chunk quality. A chunk that mixes three topics dilutes its embedding; a chunk sliced mid-sentence loses the context the model needs. The most widely used generic strategy is recursive splitting: try paragraph breaks first, then line breaks, then spaces, recursing into anything still too large — then merge small pieces with overlap. It’s the recommended generic splitter in LangChain’s RAG tutorial, and it’s short enough to own outright:
def recursive_split(text, chunk_size=500, chunk_overlap=50,
separators=None):
separators = separators or ["\n\n", "\n", " ", ""]
# first separator that actually occurs in the text
idx = next(i for i, s in enumerate(separators) if s == "" or s in text)
sep, rest = separators[idx], separators[idx + 1:]
splits = text.split(sep) if sep else list(text)
good, out = [], []
for s in splits:
if len(s) <= chunk_size:
good.append(s)
else:
if good:
out.extend(_merge(good, sep, chunk_size, chunk_overlap))
good = []
out.extend(recursive_split(s, chunk_size, chunk_overlap, rest)
if rest else [s[i:i + chunk_size]
for i in range(0, len(s), chunk_size)])
if good:
out.extend(_merge(good, sep, chunk_size, chunk_overlap))
return out
def _merge(pieces, sep, chunk_size, chunk_overlap):
docs, cur, total = [], [], 0
for p in pieces:
if cur and total + len(sep) + len(p) > chunk_size:
docs.append(sep.join(cur))
# pop whole pieces off the front: the next chunk
# starts with ~chunk_overlap chars of repeated text
while cur and total > chunk_overlap:
total -= len(cur[0]) + (len(sep) if len(cur) > 1 else 0)
cur = cur[1:]
cur.append(p)
total += len(p) + (len(sep) if len(cur) > 1 else 0)
if cur:
docs.append(sep.join(cur))
return docs
A common starting point is a few hundred tokens per chunk with 10–20% overlap — tune from there against your eval set, not from first principles. Try it on your own text:
Two pro details the demo makes visible:
- Overlap is duplication you pay for twice — once in the index, once in every prompt that retrieves both chunks. Keep it as small as still preserves context across boundaries.
- Attach metadata to every chunk: document title, section header, date, access level. Metadata turns “search everything” into “search the right slice” — filtering by section or recency before vector search is often worth more than a fancier embedding model.
2. Hybrid retrieval: stop choosing between keywords and meaning
Dense embeddings understand paraphrase but fumble exact strings — product codes, error names, rare terms. Keyword search (BM25/TF-IDF) nails exact terms but can’t see that “reimbursement period” means “refund window.” Production systems run both and fuse the rankings.
The cleanest fusion needs no weight tuning: Reciprocal Rank Fusion. Each document scores 1 / (k + rank) in each ranking (k = 60 is the standard constant), summed across rankings. A document ranked #1 and #2 beats one ranked #4 and #1 — consensus wins, and a document appearing in only one list still contributes:
from collections import defaultdict
def rrf_fuse(rankings: list[list[str]], k: int = 60):
scores = defaultdict(float)
for ranking in rankings:
for rank, doc_id in enumerate(ranking, start=1): # 1-based ranks
scores[doc_id] += 1.0 / (k + rank)
return sorted(scores.items(), key=lambda kv: kv[1], reverse=True)
fused = rrf_fuse([
["doc_refund", "doc_pricing", "doc_uptime"], # sparse (BM25) ranking
["doc_support", "doc_refund", "doc_uptime"], # dense ranking
])
# → doc_refund wins: 1/61 + 1/62 beats doc_support's 1/64 + 1/61
3. Rerank the shortlist with a cross-encoder
Vector search is a bi-encoder: it embeds the query and each chunk independently, then compares vectors. Fast, but it never lets the query and chunk actually read each other. A cross-encoder feeds the (query, chunk) pair through the model jointly and outputs a relevance score — slower, but markedly more precise. The documented pattern from sentence-transformers is a two-stage pipeline: bi-encoder retrieves the top 20–50, cross-encoder reranks them down to the top 5 that reach the prompt:
from sentence_transformers import CrossEncoder
reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
candidates = [chunk for chunk, _ in fused[:20]] # from hybrid retrieval
scores = reranker.predict([[query, c] for c in candidates])
top5 = [candidates[i] for i in scores.argsort()[::-1][:5]]
CrossEncoder(...).predict() on (query, document) pairs and the .rank() convenience method are straight from the sentence-transformers docs. Because the reranker only ever sees a few dozen candidates, its cost stays bounded even on a large corpus.
4. Rewrite the query before you search it
Users don’t write search queries — they write questions, often vague ones. Three cheap transformations, applied before retrieval:
- Multi-query: ask the LLM to generate 3–4 paraphrases of the question, retrieve for each, fuse with RRF. Catches the phrasing your corpus actually uses.
- HyDE (Hypothetical Document Embeddings): ask the LLM to write a hypothetical answer, then search with the embedding of that answer instead of the question. Answers resemble documents more than questions do, so the similarity search lands closer.
- Step-back: for complex questions, first ask the broader question (“what is the refund policy?”) to retrieve context, then answer the specific one.
All three trade extra LLM calls for better recall. Measure whether the gain survives your eval set before paying for it on every query.
5. The eval harness: what makes it shippable
Everything above is tuning. Evaluation is what tells you what to tune. A minimal harness runs on every deploy and checks two things:
- Retrieval: for each labelled Q&A pair, is a known-good chunk in the top-k? (hit-rate@k)
- Faithfulness: is every claim in the answer supported by the retrieved context? (LLM-as-judge, or a lexical proxy in CI)
def hit_rate(cases, retrieve_fn, k=3) -> float:
hits = sum(1 for c in cases
if c.must_retrieve in [cid for cid, _ in retrieve_fn(c.question, k)])
return hits / len(cases)
The fuller harness from this article’s research code adds a faithfulness judge that flags answer sentences whose content words don’t appear in the context — on a test set it passed two good answers and caught a deliberately hallucinated phone number. In production, replace the lexical proxy with an LLM judge and track the established metric families: faithfulness (claims supported by context), answer relevancy, context precision/recall (did we retrieve what’s needed, and only what’s needed). The RAGAS project standardized these names; whatever library you use, those four questions are the ones to answer.
The debugging workflow this enables is the real payoff: when the harness fails, the pattern of failure points at the stage — retrieval misses mean fix chunking or the retriever; faithful-but-wrong answers mean fix the prompt; slow queries mean check your index. Without the harness you’re tuning blind.
6. Production concerns, briefly
- Index choice. Brute-force search is fine to thousands of chunks. Past that, approximate indexes trade a little recall for a lot of speed: Faiss — “a library for efficient similarity search of dense vectors” — offers exact
IndexFlatL2baselines, graph-based HNSW, and compressed IVF/PQ indexes that fit billions of vectors on one server, with GPU indexes as drop-in replacements. Start exact; move to approximate when latency says so. - Updates. Re-embed only changed documents, and version your index alongside the embedding model — swapping models without re-indexing silently breaks similarity.
- Latency & cost budgets. Track tokens and milliseconds per query at p95. Query rewriting, reranking, and judge calls each add latency; the eval harness tells you which ones earn their keep.
- Citations. Keep the
[Source N]labels from the starter through generation, and render them as links. A grounded answer the user can’t verify is only half trustworthy. - Security. Treat retrieved chunks as untrusted data. A malicious document in your corpus can carry instructions the model will follow — indirect prompt injection. LangChain’s RAG tutorial is explicit that no prompt or delimiter strategy fully prevents it: validate outputs, scope retrieval permissions per user, and never let retrieved text drive privileged actions.
The shape of a mature system
Naive RAG: chunk → embed → search → pray. Production RAG: structure-aware chunking with metadata → rewritten queries → hybrid retrieval fused with RRF → cross-encoder rerank → cited generation → an eval harness grading every deploy and pointing at the next bottleneck. Each stage is independently testable, which is what makes the whole thing operable.
The starter article’s script is the skeleton. Everything in this article is muscle you add where the harness says it hurts.
Sources & further reading
- LangChain RAG tutorial — indexing pipeline,
RecursiveCharacterTextSplitterrecommendation, and the prompt-injection security notes - Sentence Transformers quickstart — bi-encoder / cross-encoder two-stage retrieval,
CrossEncoder.predictand.rank - Faiss (Meta) —
IndexFlatL2, HNSW/NSG graph indexes, IVF/PQ compression, GPU drop-in indexes - Reciprocal Rank Fusion (Cormack, Clarke & Buettcher, SIGIR 2009) — the
1/(k+rank)fusion formula - RAGAS — the standard faithfulness / answer-relevancy / context-precision metric family for RAG evaluation