GraphRAG: How AI Answers Questions Hidden Across Many Documents
Standard Retrieval Augmented Generation (RAG) operates through a straightforward pipeline where documents are split into token chunks and transformed into vectors via an embedding model. These vectors are stored in a vector index to facilitate proximity-based semantic search against incoming query vectors. During query execution, the closest document chunks are retrieved and included in the language model prompt to generate grounded answers. The entire workflow relies on the core assumption that queries and their corresponding answers share lexical or semantic resemblance. The assumption that questions resemble their answers holds for local queries but breaks down for global queries that require reasoning across an entire corpus. For global queries, standard vector retrieval surfaces documents with superficial vocabulary overlap rather than synthesizing information distributed across hundreds of files. Microsoft's evaluations show that increasing context windows up to 64,000 tokens does not bridge this gap in comprehensiveness, diversity, and source quality compared to GraphRAG. What is commonly diagnosed as model hallucination is often actually a retrieval issue where fluent text is generated from irrelevant retrieved context. Knowledge graphs enhance retrieval and answer quality by modeling relationships between entities across documents rather than treating text chunks independently. By aggregating extracted entities and typed relations from multiple sources, knowledge graphs create multi-hop paths connecting disparate information that no single document holds. Real-world implementations, such as LinkedIn's support ticket system presented at SIGIR 2024, demonstrate significant production gains including higher retrieval rank and reduced resolution times. Modern GraphRAG architectures typically combine lexical graphs of document structures with entity graphs of domain concepts to query across both.
閱讀原文 ↗目錄
Retrieval Basics
Standard Retrieval Augmented Generation (RAG) operates through a straightforward pipeline where documents are split into token chunks and transformed into vectors via an embedding model. These vectors are stored in a vector index to facilitate proximity-based semantic search against incoming query vectors. During query execution, the closest document chunks are retrieved and included in the language model prompt to generate grounded answers. The entire workflow relies on the core assumption that queries and their corresponding answers share lexical or semantic resemblance.
- Standard RAG partitions documents into chunks ranging from a few hundred to a few thousand tokens.
- An embedding model maps text chunks into numeric vectors representing semantic meaning within a shared vector space.
- At query time, the vector index identifies and returns chunk vectors that are numerically closest to the question vector.
- Retrieved chunk text is injected into the prompt alongside the question for the language model to generate an answer.
- The RAG architecture assumes that text providing an answer semantically or lexically resembles the query itself.
Similarity Limits
The assumption that questions resemble their answers holds for local queries but breaks down for global queries that require reasoning across an entire corpus. For global queries, standard vector retrieval surfaces documents with superficial vocabulary overlap rather than synthesizing information distributed across hundreds of files. Microsoft's evaluations show that increasing context windows up to 64,000 tokens does not bridge this gap in comprehensiveness, diversity, and source quality compared to GraphRAG. What is commonly diagnosed as model hallucination is often actually a retrieval issue where fluent text is generated from irrelevant retrieved context.
- Microsoft's GraphRAG documentation divides queries into local queries focusing on specific text regions and global queries requiring corpus-wide reasoning.
- Vector retrieval fails on global queries because nearest-neighbor search matches coincidental phrasing rather than aggregating distributed facts.
- Increasing context windows from 8,000 to 64,000 tokens fails to match GraphRAG's performance on global questions.
- Vector search with large context windows still falls short on comprehensiveness, diversity, and source material quality.
- Failures attributed to model hallucination often stem from retrieval returning irrelevant material that the model subsequently generates fluent responses from.
Knowledge Graphs
Knowledge graphs enhance retrieval and answer quality by modeling relationships between entities across documents rather than treating text chunks independently. By aggregating extracted entities and typed relations from multiple sources, knowledge graphs create multi-hop paths connecting disparate information that no single document holds. Real-world implementations, such as LinkedIn's support ticket system presented at SIGIR 2024, demonstrate significant production gains including higher retrieval rank and reduced resolution times. Modern GraphRAG architectures typically combine lexical graphs of document structures with entity graphs of domain concepts to query across both.
- Knowledge graphs store corpora as entities and typed relationships, both carrying plain-text descriptions.
- Aggregating entities and edges across multiple documents reveals paths and connections that no single document contains.
- LinkedIn's customer service team implemented a knowledge graph for support tickets, improving mean reciprocal rank by 77.6 percent and reducing median per-issue resolution time by 28.6 percent in production.
- LinkedIn published their knowledge graph retrieval findings at SIGIR in 2024.
- Neo4j distinguishes between lexical graphs (linking documents to chunks) and entity graphs (linking described entities).
- Most GraphRAG systems construct both lexical and entity graphs and query across them.
Graph Construction
Constructing a knowledge graph from raw documents requires a multi-phase pipeline where the majority of computation and expense is concentrated in graph extraction and merging. Microsoft's indexing workflow involves text chunking, LLM extraction of entities and relationships, merging and summarizing redundant descriptions, community clustering, and embedding into a vector store. The two LLM passes across a corpus represent approximately 75 percent of total indexing costs. To mitigate this expense, alternatives like FastGraphRAG substitute LLMs with traditional NLP techniques, though this trades cost efficiency for increased noise.
- Microsoft's graph indexing workflow consists of six phases, including chunking, LLM extraction, merging/compression, optional claim extraction, community clustering, and report embedding.
- The merge step is computationally expensive because entities appearing across hundreds of text units require separate LLM passes to reconcile into coherent descriptions.
- Every extracted entity, relationship, and claim maintains provenance via pointers back to source text units for paragraph-level citations.
- Microsoft estimates that graph extraction accounts for approximately 75 percent of total indexing costs.
- FastGraphRAG reduces indexing costs by swapping LLMs for traditional NLP, extracting noun phrases and co-occurrences at the expense of higher graph noise.
Community Detection
GraphRAG enables answering whole-collection questions by executing hierarchical Leiden clustering across an entity graph to partition it into multi-level communities. For each community across every level, a language model pre-generates a community report detailing key entities, relationships, and claims during the indexing phase. Choosing the appropriate hierarchy level determines the balance between answer thoroughness and operational costs in time and tokens.
- Hierarchical Leiden clustering recursively partitions entity graphs into multi-level communities until reaching a minimum size threshold.
- Higher levels (such as Level 0) provide broad coverage over large regions, while deeper levels partition into narrower, fine-grained sub-communities.
- Community reports containing summaries, key entities, relationships, and claims are generated during indexing before queries occur.
- Pre-generated summaries allow global, whole-collection questions to be answered directly from pre-existing text.
- Selecting lower hierarchical levels produces more detailed answers but increases token consumption and query latency due to processing a larger volume of reports.
Query Modes
GraphRAG supports multiple retrieval query modes tailored to different question types across a knowledge graph and report hierarchy. Local search identifies specific entry-point entities and expands along five parallel directions—text units, community reports, neighbors, relationships, and covariates—before ranking and packing them into a context window. Global search operates over community reports using a map-reduce process to answer broad, corpus-wide questions without relying on specific entity lookups. Additionally, GraphRAG provides DRIFT search, which combines community reports and follow-up local searches, as well as a basic search mode using standard top-k vector retrieval.
- Local search identifies entry-point entities via embeddings and expands across five parallel directions: text units, community reports, neighboring entities, relationships, and covariates.
- Local search functions as a bounded, structured gather-and-rank operation rather than open-ended pathfinding.
- Global search bypasses the entity graph, processing shuffled community report batches through a map-reduce pipeline with intermediate numerical importance ratings.
- DRIFT search blends global and local approaches by producing an initial answer from community reports, generating follow-up questions, and running local search on them.
- GraphRAG includes a basic search mode that uses standard top-k vector retrieval.
Cost Tradeoffs
GraphRAG significantly redistributes retrieval-augmented generation costs compared to standard RAG, incurring heavy indexing expenses due to multiple language model passes and community summarization. Updating the index with new documents creates an ongoing operational commitment as hierarchies shift and material is re-summarized. Microsoft developed LazyGraphRAG to address this by replacing LLM-based indexing with traditional NLP, which reduces indexing cost to match vector RAG and lowers query cost by over 700x while maintaining global-query quality. Nevertheless, full GraphRAG remains valuable for its human-readable community reports and excels in comprehensiveness and diversity, though its faithfulness matches standard RAG rather than eliminating hallucinations.
- Standard vector RAG requires only one embedding pass at index time and a single nearest-neighbor lookup per query.
- GraphRAG indexing requires two language model passes over the corpus and report generation across community hierarchy levels.
- LazyGraphRAG reduces indexing cost to 0.1 percent of full GraphRAG by utilizing NLP instead of LLMs and deferring LLM tasks to query time.
- LazyGraphRAG cuts query costs by more than a factor of 700 while maintaining comparable global-query quality.
- Microsoft notes that pre-built entity, relationship, and community summaries provide utility beyond question answering because people read and share them directly.
- Evaluations show vector RAG outperforms GraphRAG for localized queries where answers reside in specific text regions.
- GraphRAG provides advantages in comprehensiveness and diversity, but scores similarly to baseline RAG on faithfulness and factual accuracy.
Agentic Retrieval
Agentic RAG addresses the limitations of fixed retrieval pipelines by using a language model to dynamically classify queries and select tailored retrieval strategies, such as vector search, SQL, or web search. LlamaIndex illustrates a two-layer routing design where a composite retriever selects an index and an auto-routed mode chooses the retrieval method within that index. While this flexibility improves accuracy across diverse questions, it introduces latency, higher per-query costs, and debugging hurdles from routing errors. The paradigm marks an evolutionary step beyond basic RAG, advanced RAG, and GraphRAG.
- Agentic RAG uses an LLM to classify incoming queries, select a specific retrieval strategy, execute it, and synthesize results.
- Supported retrieval options can span vector search, corpus-wide global search, SQL queries, and web search.
- LlamaIndex implements a two-layer routing pattern using a composite retriever to pick an index and an auto-routed mode to select the method within that index.
- Pre-retrieval LLM calls introduce additional latency and increased financial cost per query.
- Routing failures complicate debugging because an appropriate retrieval method can execute poorly if misassigned to a query.
- Agentic RAG represents the top of a performance ladder evolving from basic RAG, advanced RAG, and GraphRAG.
Conclusion
GraphRAG addresses the shortcomings of standard vector retrieval for global queries by extracting a knowledge graph and generating hierarchical community summaries across a corpus. Graph extraction is computationally expensive, driving around 75 percent of total indexing costs, which has motivated lower-cost alternatives like LazyGraphRAG. While local search expands from matched graph entities and vector RAG remains better suited for local queries, global search performs map-reduce operations over pre-generated community reports. Consequently, agentic retrieval frameworks can dynamically pick between vector and graph-based strategies depending on query needs.
- Local queries target localized text regions, whereas global queries require synthesizing information across large parts of a collection.
- Standard similarity search and expanded context windows (up to 64,000 tokens) still struggle with global questions compared to GraphRAG.
- Graph extraction accounts for approximately 75 percent of GraphRAG's indexing cost.
- Hierarchical Leiden clustering partitions the entity graph into multi-resolution communities with pre-generated community reports.
- LazyGraphRAG shifts language model computation to query time, reducing indexing costs to 0.1 percent of standard GraphRAG.
- Agentic retrieval dynamically routes queries between vector RAG and GraphRAG based on whether the query is local or global.