RAG in Production: Chunking, Hybrid Search, Reranking, and Evals

Naive RAG fails silently — the answer looks fine while the retrieval missed. Here's the production upgrade path: recursive chunking, hybrid search with RRF fusion, cross-encoder reranking, query rewriting, and an eval harness that tells you which stage to fix. Core algorithms executed and verified.

The starter system from RAG from Scratch works — on the demo questions. Ship it, and you’ll meet its failure mode: it fails silently. The answer reads fluently while the retrieval quietly missed, and the model filled the gap with something plausible. Nobody notices until a customer does.

The production upgrade isn’t one trick. It’s a loop: better chunking, better retrieval, reranking, and — the part most teams skip — an eval harness that tells you which stage to fix.

Production RAG loop: query rewriting, hybrid retrieval, reranking, generation with citations, and an eval harness scoring every answer and feeding failures back into chunking and retrieval tuning
The full loop. Every stage is tunable; the eval harness decides what to tune. Click to expand.

1. Chunking: the highest-leverage boring work

Retrieval quality is bounded by chunk quality. A chunk that mixes three topics dilutes its embedding; a chunk sliced mid-sentence loses the context the model needs. The most widely used generic strategy is recursive splitting: try paragraph breaks first, then line breaks, then spaces, recursing into anything still too large — then merge small pieces with overlap. It’s the recommended generic splitter in LangChain’s RAG tutorial, and it’s short enough to own outright:

def recursive_split(text, chunk_size=500, chunk_overlap=50,
                    separators=None):
    separators = separators or ["\n\n", "\n", " ", ""]
    # first separator that actually occurs in the text
    idx = next(i for i, s in enumerate(separators) if s == "" or s in text)
    sep, rest = separators[idx], separators[idx + 1:]
    splits = text.split(sep) if sep else list(text)

    good, out = [], []
    for s in splits:
        if len(s) <= chunk_size:
            good.append(s)
        else:
            if good:
                out.extend(_merge(good, sep, chunk_size, chunk_overlap))
                good = []
            out.extend(recursive_split(s, chunk_size, chunk_overlap, rest)
                       if rest else [s[i:i + chunk_size]
                                     for i in range(0, len(s), chunk_size)])
    if good:
        out.extend(_merge(good, sep, chunk_size, chunk_overlap))
    return out

def _merge(pieces, sep, chunk_size, chunk_overlap):
    docs, cur, total = [], [], 0
    for p in pieces:
        if cur and total + len(sep) + len(p) > chunk_size:
            docs.append(sep.join(cur))
            # pop whole pieces off the front: the next chunk
            # starts with ~chunk_overlap chars of repeated text
            while cur and total > chunk_overlap:
                total -= len(cur[0]) + (len(sep) if len(cur) > 1 else 0)
                cur = cur[1:]
        cur.append(p)
        total += len(p) + (len(sep) if len(cur) > 1 else 0)
    if cur:
        docs.append(sep.join(cur))
    return docs

A common starting point is a few hundred tokens per chunk with 10–20% overlap — tune from there against your eval set, not from first principles. Try it on your own text:

LIVE DEMO — CHUNKING LAB
4 chunksavg 226 charsoverlap duplication 14%
CHUNK 1232 chars
Retrieval augmented generation grounds large language model answers in your own documents. Instead of hoping the model memorized your docs during training, you retrieve the relevant passages at query time and put them in the prompt.
CHUNK 2300 chars
query time and put them in the prompt. The indexing pipeline has four steps: load, split, embed, store. Chunking is the step everyone underestimates. A chunk that is too large dilutes the embedding signal with unrelated text. A chunk that is too small loses the context the model needs to answer. A
CHUNK 3117 chars
he context the model needs to answer. A common starting point is a few hundred tokens with 10 to 20 percent overlap.
CHUNK 4255 chars
with 10 to 20 percent overlap. At query time you embed the question with the same model, run a similarity search over the chunk vectors, and stuff the top-k chunks into the prompt. The model then answers from the retrieved context instead of its weights.
Highlighted text is carried over from the previous chunk. Switch to fixed-size and watch it slice mid-sentence — then compare with recursive splitting at the same size.

Two pro details the demo makes visible:

  • Overlap is duplication you pay for twice — once in the index, once in every prompt that retrieves both chunks. Keep it as small as still preserves context across boundaries.
  • Attach metadata to every chunk: document title, section header, date, access level. Metadata turns “search everything” into “search the right slice” — filtering by section or recency before vector search is often worth more than a fancier embedding model.

2. Hybrid retrieval: stop choosing between keywords and meaning

Dense embeddings understand paraphrase but fumble exact strings — product codes, error names, rare terms. Keyword search (BM25/TF-IDF) nails exact terms but can’t see that “reimbursement period” means “refund window.” Production systems run both and fuse the rankings.

The cleanest fusion needs no weight tuning: Reciprocal Rank Fusion. Each document scores 1 / (k + rank) in each ranking (k = 60 is the standard constant), summed across rankings. A document ranked #1 and #2 beats one ranked #4 and #1 — consensus wins, and a document appearing in only one list still contributes:

from collections import defaultdict

def rrf_fuse(rankings: list[list[str]], k: int = 60):
    scores = defaultdict(float)
    for ranking in rankings:
        for rank, doc_id in enumerate(ranking, start=1):  # 1-based ranks
            scores[doc_id] += 1.0 / (k + rank)
    return sorted(scores.items(), key=lambda kv: kv[1], reverse=True)

fused = rrf_fuse([
    ["doc_refund", "doc_pricing", "doc_uptime"],   # sparse (BM25) ranking
    ["doc_support", "doc_refund", "doc_uptime"],   # dense ranking
])
# → doc_refund wins: 1/61 + 1/62 beats doc_support's 1/64 + 1/61
Hybrid retrieval: the query fans out to sparse BM25 keyword search and dense vector search in parallel; both ranked lists fuse with reciprocal rank fusion, then a cross-encoder reranks the fused candidates into the final top-k
Two retrievers, complementary blind spots. RRF fuses; the reranker refines. Click to expand.

3. Rerank the shortlist with a cross-encoder

Vector search is a bi-encoder: it embeds the query and each chunk independently, then compares vectors. Fast, but it never lets the query and chunk actually read each other. A cross-encoder feeds the (query, chunk) pair through the model jointly and outputs a relevance score — slower, but markedly more precise. The documented pattern from sentence-transformers is a two-stage pipeline: bi-encoder retrieves the top 20–50, cross-encoder reranks them down to the top 5 that reach the prompt:

from sentence_transformers import CrossEncoder

reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2")
candidates = [chunk for chunk, _ in fused[:20]]          # from hybrid retrieval
scores = reranker.predict([[query, c] for c in candidates])
top5 = [candidates[i] for i in scores.argsort()[::-1][:5]]

CrossEncoder(...).predict() on (query, document) pairs and the .rank() convenience method are straight from the sentence-transformers docs. Because the reranker only ever sees a few dozen candidates, its cost stays bounded even on a large corpus.

4. Rewrite the query before you search it

Users don’t write search queries — they write questions, often vague ones. Three cheap transformations, applied before retrieval:

  • Multi-query: ask the LLM to generate 3–4 paraphrases of the question, retrieve for each, fuse with RRF. Catches the phrasing your corpus actually uses.
  • HyDE (Hypothetical Document Embeddings): ask the LLM to write a hypothetical answer, then search with the embedding of that answer instead of the question. Answers resemble documents more than questions do, so the similarity search lands closer.
  • Step-back: for complex questions, first ask the broader question (“what is the refund policy?”) to retrieve context, then answer the specific one.

All three trade extra LLM calls for better recall. Measure whether the gain survives your eval set before paying for it on every query.

5. The eval harness: what makes it shippable

Everything above is tuning. Evaluation is what tells you what to tune. A minimal harness runs on every deploy and checks two things:

  1. Retrieval: for each labelled Q&A pair, is a known-good chunk in the top-k? (hit-rate@k)
  2. Faithfulness: is every claim in the answer supported by the retrieved context? (LLM-as-judge, or a lexical proxy in CI)
def hit_rate(cases, retrieve_fn, k=3) -> float:
    hits = sum(1 for c in cases
               if c.must_retrieve in [cid for cid, _ in retrieve_fn(c.question, k)])
    return hits / len(cases)

The fuller harness from this article’s research code adds a faithfulness judge that flags answer sentences whose content words don’t appear in the context — on a test set it passed two good answers and caught a deliberately hallucinated phone number. In production, replace the lexical proxy with an LLM judge and track the established metric families: faithfulness (claims supported by context), answer relevancy, context precision/recall (did we retrieve what’s needed, and only what’s needed). The RAGAS project standardized these names; whatever library you use, those four questions are the ones to answer.

The debugging workflow this enables is the real payoff: when the harness fails, the pattern of failure points at the stage — retrieval misses mean fix chunking or the retriever; faithful-but-wrong answers mean fix the prompt; slow queries mean check your index. Without the harness you’re tuning blind.

6. Production concerns, briefly

  • Index choice. Brute-force search is fine to thousands of chunks. Past that, approximate indexes trade a little recall for a lot of speed: Faiss — “a library for efficient similarity search of dense vectors” — offers exact IndexFlatL2 baselines, graph-based HNSW, and compressed IVF/PQ indexes that fit billions of vectors on one server, with GPU indexes as drop-in replacements. Start exact; move to approximate when latency says so.
  • Updates. Re-embed only changed documents, and version your index alongside the embedding model — swapping models without re-indexing silently breaks similarity.
  • Latency & cost budgets. Track tokens and milliseconds per query at p95. Query rewriting, reranking, and judge calls each add latency; the eval harness tells you which ones earn their keep.
  • Citations. Keep the [Source N] labels from the starter through generation, and render them as links. A grounded answer the user can’t verify is only half trustworthy.
  • Security. Treat retrieved chunks as untrusted data. A malicious document in your corpus can carry instructions the model will follow — indirect prompt injection. LangChain’s RAG tutorial is explicit that no prompt or delimiter strategy fully prevents it: validate outputs, scope retrieval permissions per user, and never let retrieved text drive privileged actions.

The shape of a mature system

Naive RAG: chunk → embed → search → pray. Production RAG: structure-aware chunking with metadata → rewritten queries → hybrid retrieval fused with RRF → cross-encoder rerank → cited generation → an eval harness grading every deploy and pointing at the next bottleneck. Each stage is independently testable, which is what makes the whole thing operable.

The starter article’s script is the skeleton. Everything in this article is muscle you add where the harness says it hurts.

Sources & further reading

  • LangChain RAG tutorial — indexing pipeline, RecursiveCharacterTextSplitter recommendation, and the prompt-injection security notes
  • Sentence Transformers quickstart — bi-encoder / cross-encoder two-stage retrieval, CrossEncoder.predict and .rank
  • Faiss (Meta) — IndexFlatL2, HNSW/NSG graph indexes, IVF/PQ compression, GPU drop-in indexes
  • Reciprocal Rank Fusion (Cormack, Clarke & Buettcher, SIGIR 2009) — the 1/(k+rank) fusion formula
  • RAGAS — the standard faithfulness / answer-relevancy / context-precision metric family for RAG evaluation

Keep reading