How LLMs Can Find a Needle in a Haystack
Large language models lack internal access to private corporate documents and require external evidence provided in context to answer questions accurately. Retrieval-augmented generation (RAG) addresses this by using embedding models and vector databases to retrieve relevant document passages before passing them to the LLM. The quality of the final generated answer depends heavily on the retrieval stage, as incomplete or outdated retrieved context leads to inaccurate answers despite the model's capabilities. Consequently, document collections must be indexed and searchable at the right granularity to support reliable outputs. Treating an entire document as a single searchable unit creates an overly broad representation that obscures specific rules. To resolve this, retrieval applications divide documents into smaller units called chunks to pinpoint individual passages. Chunking requires balancing precision and context, as undersized chunks may omit critical qualifications while oversized chunks add irrelevant noise. Techniques like including section headings and providing limited overlap between neighboring chunks help retain necessary context so passages remain understandable in isolation. Embeddings enable semantic search across differing wordings by representing passages as high-dimensional numerical vectors in a shared space. Meaning is encoded not in isolated dimensions, but through holistic patterns across vectors where semantically related passages cluster closely together. Effective retrieval requires queries and document chunks to be mapped into the same space using compatible encoders rather than models that merely share vector lengths. Additionally, vector records must retain connections to the underlying passage text, chunk identifiers, and metadata so language models can interpret and maintain the retrieved information.
閱讀原文 ↗目錄
- 01The LLM Needs Evidence to Answer Questions
- 02How Documents are Turned into Searchable Passages
- 03How LLMs Find the Meaning Behind Different Words
- 04How Close Is Close Enough?
- 05Why Searching Every Passage Is Too Expensive
- 06Following Connections to the Right Neighborhood
- 07How Much Searching Is Enough?
- 08The Closest Match Might Be the Wrong Policy
- 09What Happens When the Answer Changes?
- 10From Promising Matches to a Supported Answer
- 11Conclusion
The LLM Needs Evidence to Answer Questions
Large language models lack internal access to private corporate documents and require external evidence provided in context to answer questions accurately. Retrieval-augmented generation (RAG) addresses this by using embedding models and vector databases to retrieve relevant document passages before passing them to the LLM. The quality of the final generated answer depends heavily on the retrieval stage, as incomplete or outdated retrieved context leads to inaccurate answers despite the model's capabilities. Consequently, document collections must be indexed and searchable at the right granularity to support reliable outputs.
- Language models do not inherently know the contents of private documents and must be supplied with relevant evidence at query time.
- Retrieval-augmented generation (RAG) is a pattern where an application retrieves relevant text passages and feeds them to an LLM to answer questions and cite sources.
- Embedding models convert questions into numerical representations, which vector databases store and search to identify relevant passages.
- The accuracy of the generated answer depends directly on the pre-generation retrieval stage; incomplete or outdated retrieved context leads to faulty answers.
- Effective RAG systems require document collections to be searchable at an appropriate level of detail.
How Documents are Turned into Searchable Passages
Treating an entire document as a single searchable unit creates an overly broad representation that obscures specific rules. To resolve this, retrieval applications divide documents into smaller units called chunks to pinpoint individual passages. Chunking requires balancing precision and context, as undersized chunks may omit critical qualifications while oversized chunks add irrelevant noise. Techniques like including section headings and providing limited overlap between neighboring chunks help retain necessary context so passages remain understandable in isolation.
- Indexing an entire document as a single unit creates a broad representation that obscures specific policy rules.
- Chunking partitions documents into smaller units, enabling search systems to return precise passages instead of entire source documents.
- Chunk size involves a trade-off between search precision and contextual completeness.
- Separating a rule from its condition or exception across different chunks can lead an assistant to generate incorrect answers.
- Preserving section headings and using limited overlap between adjacent chunks helps maintain standalone comprehension.
How LLMs Find the Meaning Behind Different Words
Embeddings enable semantic search across differing wordings by representing passages as high-dimensional numerical vectors in a shared space. Meaning is encoded not in isolated dimensions, but through holistic patterns across vectors where semantically related passages cluster closely together. Effective retrieval requires queries and document chunks to be mapped into the same space using compatible encoders rather than models that merely share vector lengths. Additionally, vector records must retain connections to the underlying passage text, chunk identifiers, and metadata so language models can interpret and maintain the retrieved information.
- Embedding models transform text passages into vectors containing hundreds or thousands of decimal dimensions.
- Semantic meaning is distributed across the complete vector pattern rather than assigned to individual labeled dimensions.
- Queries and document chunks must be embedded into the same vector space using compatible encoders.
- Having matching vector dimensions between two models does not imply encoder compatibility.
- Searchable records must link vector embeddings to raw text, chunk IDs, and metadata like version and section for LLM interpretation.
How Close Is Close Enough?
Search systems require similarity or distance metrics to define proximity between vector representations. Core metrics such as cosine similarity, Euclidean distance, and dot product behave differently with respect to vector magnitude, but normalization to unit length makes their rankings equivalent. Selecting an appropriate metric depends directly on the embedding model's intended design rather than popularity. Furthermore, similarity scores represent geometric relationships between vectors rather than the probability that a retrieved passage is factually correct.
- Cosine similarity compares vector directions while ignoring vector length.
- Euclidean distance measures straight-line distance between endpoints, making it sensitive to vector length.
- Dot product reflects both directional alignment and vector length.
- When vectors are normalized to unit length, dot product equals cosine similarity and Euclidean distance yields the identical ranking.
- The choice of distance metric should strictly align with the embedding model's design.
- Similarity scores quantify vector relationships and do not represent the probability of an answer being correct.
Why Searching Every Passage Is Too Expensive
Flat vector search compares queries against every eligible vector to yield exact nearest neighbors, but its computational cost scales linearly with collection size and vector dimensionality. To reduce search expense, Inverted-File (IVF) indexing organizes vectors into clusters around representative centers, querying only a selected subset of groups. The nprobe parameter dictates how many groups are searched, trading off computational speed for recall. As a result, IVF introduces approximation by accepting the risk of bypassing relevant passages located in skipped groups.
- Flat index search comparison scales linearly (O(n)) with collection size and increases with vector dimensionality.
- Flat search guarantees exact nearest neighbors numerically, though it does not ensure semantic correctness.
- Inverted-File (IVF) indexing reduces comparison work by grouping vectors around cluster centers and searching only selected groups.
- The nprobe parameter controls how many vector groups are examined during an IVF search.
- Searching more groups increases recall but also increases computational work.
- IVF is an approximate search method because relevant passages in skipped groups may be missed.
Following Connections to the Right Neighborhood
Hierarchical Navigable Small World (HNSW) is a graph-based indexing method designed for approximate nearest-neighbor search that avoids exhaustive vector comparisons. It structures vector links across hierarchical layers, navigating from sparse upper layers down to a fully populated bottom layer to refine searches efficiently. While HNSW significantly reduces comparison operations, routing shortcuts can occasionally miss true nearest neighbors. Choosing between HNSW, flat search, and Inverted-File indexing depends on factors like available memory, query volume, and latency requirements rather than vector count alone.
- HNSW organizes vector connections into a multi-layered graph with sparse upper layers and a complete bottom layer.
- Search navigation operates top-down, moving quickly toward the target neighborhood before conducting finer-grained candidate exploration at the base.
- Like Inverted-File indexing, HNSW performs approximate nearest-neighbor search, reducing comparison costs at the potential expense of missing true nearest neighbors.
- No fixed collection size dictates when flat search should be replaced by Inverted-File or HNSW indexing.
- System selection depends heavily on memory availability, vector dimensions, query volume, filtering, and latency constraints.
How Much Searching Is Enough?
Approximate search introduces an inherent tradeoff between retrieval speed and nearest-neighbor recall. Evaluating systems requires distinguishing between index recall and actual evidence relevance to user queries. In the HNSW algorithm, parameters such as M, ef_construction, and ef_search govern graph connectivity, construction cost, and the breadth of candidate exploration. Tuning these settings must be driven by empirical testing on representative questions rather than simply increasing latency without answering accuracy gains.
- Approximate search creates a measurable tradeoff between search speed and recall compared to exact nearest-neighbor search.
- Index recall measures fidelity to exact search results, not whether the retrieved passages actually answer a user's question.
- In HNSW, M dictates graph connectivity, trading higher memory use and build time for improved recall.
- ef_construction determines search breadth during insertion, while ef_search controls candidate exploration breadth during queries.
- The candidate exploration breadth (ef_search) is independent of the final count of returned results.
- Parameter tuning should be validated against representative questions to ensure increased latency provides real retrieval benefits.
The Closest Match Might Be the Wrong Policy
Similarity search alone can retrieve irrelevant documents if eligibility criteria like region or effective date are not enforced. Metadata filtering addresses this by restricting results using approaches such as pre-filtering, which evaluates criteria before ranking, or post-filtering, which filters after candidate retrieval. In vector and graph search, neither naive pre-filtering nor post-filtering is universally optimal because disallowed nodes can still provide critical traversal paths. Consequently, search engines often integrate filtering directly into search via filtering-aware graph structures, specialized traversal strategies, or exact scans for small subsets.
- Metadata filtering is necessary when similarity ranking alone would surface semantically close but ineligible records.
- Pre-filtering identifies eligible records prior to similarity ranking, whereas post-filtering prunes records after candidate retrieval.
- Post-filtering can fail to return enough valid results if top similarity candidates predominantly fail eligibility criteria.
- In graph search, blocking ineligible nodes completely during traversal can make eligible neighbor nodes unreachable.
- Filtering can be integrated directly into search algorithms using filtering-aware graphs, traversal modifications, or exact scans.
What Happens When the Answer Changes?
Document collections must be maintained carefully when source information changes to prevent retrieval systems from delivering contradictory or outdated answers. Modifying embedded text requires generating new embeddings, whereas metadata updates can often be executed without re-embedding if the metadata was not vectorized. Coordinating updates using version identifiers and active-version rules helps avoid temporary gaps or duplication during ingestion. In addition, transitioning to a new embedding model requires full-scale planning to re-embed the entire corpus into the new model's vector space.
- Modifying source text requires new embeddings, though unchanged chunks can be reused via stable identifiers.
- Metadata-only updates can often bypass re-embedding if the metadata was not included in the embedding text.
- Deleting old chunks before inserting new ones creates temporary retrieval gaps, while inserting first risks temporary duplication.
- Version identifiers and active-version rules facilitate predictable transitions and preserve historical policy access.
- Changing an embedding model requires re-embedding the existing document collection into the new representation space.
From Promising Matches to a Supported Answer
The final retrieval phase converts candidate documents into reliable evidence for a large language model. A reranker evaluates and selects the most relevant passages from an initial set of vector search results. Hybrid search improves candidate quality by pairing semantic search using embeddings with keyword search for exact terms. System designers must also account for unanswerable queries, recognizing that finding the nearest vector match does not guarantee the existence of a factual answer.
- A reranker compares candidate passages directly against a user question to select the most relevant evidence.
- Hybrid search combines semantic retrieval via embeddings with keyword-based retrieval.
- Keyword search helps maintain exact matches for specific identifiers, names, and unusual technical terms.
- Vector search returns nearest-neighbor matches even when a database contains no relevant information for a query.
- Applications should explicitly state when retrieved policy context lacks an answer rather than hallucinating.
Conclusion
Building an effective retrieval pipeline for an LLM application requires chunking documents and representing them as vector embeddings to capture semantic similarity. As document volume expands, indexing strategies such as flat search, IVF, and HNSW navigate explicit trade-offs across query speed, memory footprint, and recall. In addition to similarity matching, systems rely on metadata filtering, continuous index hygiene, hybrid search, and reranking to select accurate and up-to-date context. Ultimately, the quality and factual reliability of the LLM's final response depend directly on this integrated retrieval workflow.
- Embeddings convert text chunks into vectors, enabling semantic retrieval across differing phrasings.
- Flat search performs exhaustive vector comparisons, whereas IVF and HNSW decrease latency via grouping and graph traversal at the expense of memory and recall trade-offs.
- Vector similarity alone is insufficient to guarantee valid evidence without metadata filtering by document type, region, or policy version.
- Regular index updates are necessary to prevent stale and duplicate documents from polluting retrieval results.
- Combining hybrid search and reranking enhances the quality of the final context supplied to the LLM.