Ask a vector-based RAG system: "Which AI labs have partnerships with cloud providers, and which models did they release?" The answer exists in your corpus, but it is scattered. One chunk says Anthropic partnered with AWS. Another, in a different document, describes the Claude models. A third covers Google DeepMind and Gemini. A fourth mentions that Mistral's founders came from DeepMind. No single chunk contains the answer, and no single chunk is "similar" to the whole question. Vector search returns a few chunks that each match part of it and misses the connections that make up the actual answer.
The limitation: chunks are islands
In a vector database, every chunk is independent. There is no record that chunk 17 and chunk 342 are about the same person, or that the company in one is a subsidiary of the company in another. Similarity search can find text that resembles the question. It cannot follow a relationship: person → founded → company → partnered with → cloud provider.
Questions that depend on relationships are common in real organisations:
- Who reports to the VP who owns the payments service?
- Which policies reference the travel policy, and would change if it changed?
- Which suppliers are connected to the delayed shipments?
- What has this customer bought, from whom, and which of those products had recalls?
The idea: store a knowledge graph too
A knowledge graph stores facts as nodes (entities: people, organisations, products, policies, concepts) and relationships (typed, directed edges: FOUNDED, PARTNERED_WITH, RELEASED, REFERENCES). Graph databases such as Neo4j store these natively and query them with languages such as Cypher, where a pattern like "organisation –PARTNERED_WITH→ organisation" is matched directly and traversed in milliseconds, however many hops away the answer is.
GraphRAG keeps the familiar vector index and adds a graph built from the same documents. At query time it uses both: vectors find the relevant passages, and the graph supplies the connected facts around them.
Building the graph with an LLM
Hand-building a knowledge graph is slow. The modern approach uses an LLM to extract it from the chunks:
- Chunk the documents as usual.
- For each chunk, ask an LLM to extract entities and relationships. LangChain's LLM Graph Transformer is one such tool. From "Anthropic, founded by former OpenAI researchers, partnered with AWS to offer Claude…" it extracts nodes (Anthropic: Organisation; OpenAI: Organisation; AWS: Organisation; Claude: Model) and relationships (Anthropic –PARTNERED_WITH→ AWS; Anthropic –RELEASED→ Claude; …).
- Constrain the schema if you want consistency: give the extractor the allowed node types (Person, Organisation, Model, Product) and relationship types. Unconstrained extraction is quick to start with but produces many near-duplicate types ("WORKS_FOR", "EMPLOYED_BY", "IS_EMPLOYEE_OF").
- Keep the source. Store each chunk as a
Documentnode linked to the entities extracted from it, so every graph fact can be traced back to its text and quoted as evidence. - Resolve entities. "Google DeepMind", "DeepMind" and "GDM" should become one node. Without entity resolution the graph fragments and traversals miss connections.
Extraction calls an LLM on every chunk, so building the graph costs far more than embedding, and extraction errors (a missed or invented relationship) become wrong facts in the graph. Review a sample, and constrain the schema for anything high-stakes.
Querying: three styles
1. Text-to-Cypher. Give an LLM the graph schema (node labels, relationship types, properties) and ask it to write a Cypher query for the question. "Which organisations partnered with AWS?" becomes a pattern match over PARTNERED_WITH relationships. The results, which are exact structured facts, go to the final LLM call for a readable answer. This is the graph equivalent of text-to-SQL from the previous module, and needs the same discipline: read-only access, validation and error-feedback retries.
2. Vector entry, graph expansion. Embed the question, find the most similar chunks or entity nodes (graph databases like Neo4j can hold vector indexes too), then expand from them along relationships, one or two hops, to collect the connected facts. This needs no query generation, and it is robust for open-ended questions.
3. Hybrid (the production default). Run vector retrieval and graph retrieval, and give the model both: the relevant passages, plus the structured facts linking the entities in them. The model uses whichever evidence the question needs. Neither alone is as good. Vectors bring rich text but no links, and graphs bring links but little nuance.
Global questions and community summaries
A different family of GraphRAG, popularised by Microsoft Research, targets global questions: "What are the main themes across all these incident reports?" No set of top-k chunks can answer that. That approach clusters the graph into communities of closely connected entities, has an LLM pre-write a summary of each community, and answers global questions by combining those summaries. It is powerful for sense-making over a whole corpus, and expensive to build.
When to use GraphRAG
GraphRAG adds a second datastore, an expensive LLM extraction step and a harder update path (re-extracting entities when documents change). Use it when the questions genuinely depend on relationships: organisational structures, supply chains, research ecosystems, interlinked regulations and policies, customer–product histories, fraud rings. For most document Q&A, where the answer sits in one or two passages, well-tuned vector or hybrid RAG is simpler, cheaper and just as good. As with every technique in this course, build a test set of your real multi-hop questions and measure whether the graph actually helps.