GraphRAG in Python: Agentic AI with Knowledge Graphs
slide 01 of 15
Welcome: Knowledge Graphs for AI
GraphRAG in Python: Agentic AI with Knowledge Graphs
Welcome. This is a deep, hands-on course on one of the most promising ideas in applied AI: Retrieval-Augmented Generation (RAG) powered by knowledge graphs. You will not just read about it — you will build a working agent that reasons over connected data, then trace exactly how it thinks.
What you will build
By the end you will have a LangChain agent backed by a Neo4j graph that answers questions no plain vector search can — questions that hop from one entity to another, like:
Where does the guy live who created the operating system that someone I know uses?
That question needs to walk a chain of relationships. A vector store retrieves similar chunks; a knowledge graph lets the agent traverse the structure.
Why this matters now
Plain RAG is powerful but has a well-documented blind spot. Microsoft’s GraphRAG paper (arXiv:2404.16130) shows that naive retrieval “struggles to connect the dots” — it fails when an answer requires combining disparate pieces of information, and it performs poorly at holistic understanding of a large corpus. Knowledge graphs address exactly this.
- Grounded answers tied to real entities and relationships, not just similar text.
- Multi-hop reasoning — follow chains of relationships, not a single lookup.
- Explainable — graph paths are easy to trace, so you can show why the agent answered as it did.
Course map
- RAG fundamentals and why grounding matters
- Knowledge graphs 101: triples, RDF, and the property graph model
- GraphRAG theory: the Microsoft two-tier architecture
- Setting up Neo4j with Docker
- Cypher, the query language, deep-dive
- Environment and Python dependencies
- Building the knowledge graph in code
- The retriever: Text2CypherRetriever
- Wiring the agent and streaming its reasoning
- A multi-hop question workshop
- Hybrid RAG: combining vector + graph
- Scaling and production
- Real-world use cases
- Anti-patterns and when NOT to use a graph
Let’s begin with the foundation: what RAG actually is, and where it runs out of road.
slide 02 of 15
RAG Fundamentals: Why Grounding
The problem with raw model weights
An LLM trained on internet-scale text knows a lot, but it has no reliable memory of your private documents, your product catalog, or your company’s internal wiki. Worse, it can confidently answer from weights that are outdated or wrong.
Retrieval-Augmented Generation
RAG changes the equation. When a user asks a question, the system:
- Retrieves relevant information from an external store (chunks, records, or graph paths)
- Augments the prompt with that retrieved context
- Generates a grounded answer from the model, constrained by the context
The result is answers that are grounded in real context you control, and that can cite their sources.
The two families of retrieval
| Retrieval | Store | Best for |
|---|---|---|
| Vector (semantic) | Embeddings + similarity search | Finding the most similar text chunks |
| Graph (structural) | Nodes + relationships | Following connections across entities |
Where plain RAG hits a wall
Vector search finds the chunks that look like the question. It struggles when the answer lives across several connected chunks — the information is there, but no single chunk contains it. That is the gap this course closes.
Baseline RAG “struggles to connect the dots” — it fails when an answer requires traversing disparate pieces of information via shared attributes. — Microsoft GraphRAG docs
Now let’s understand the substrate: what a knowledge graph actually is.
slide 03 of 15
Knowledge Graphs 101: Triples & Models
What a knowledge graph is
A knowledge graph is a knowledge base that uses a graph-structured data model — interlinked descriptions of entities (objects, events, concepts) with the relationships between them encoded as edges. (Wikipedia: “Knowledge graph”)
The fundamental unit: the triple
The atomic fact in a graph is a triple: a subject, a predicate, and an object. It reads like a sentence.
(Florian) -[:OWNS]-> (NeuralNine)
(subject) -predicate- (object)Three triples in the course’s example:
Florian OWNS NeuralNineFlorian USES LinuxLinus Torvalds CREATED Linux
Two data models
RDF (Resource Description Framework)
A W3C standard. A directed graph of triples where every part is identified by an Internationalized Resource Identifier (IRI), and the object can be a literal value. Designed for the open web and data interchange.
Property graph (what Neo4j uses)
In Neo4j’s model, a node represents an entity, is classified by labels, and holds data in properties. A relationship connects a source node to a target node, is classified by type, and can hold its own properties. This is the model Cypher queries.
Why graphs beat bolted-on linking
Connecting pieces of information in non-graph systems means storing a lot of extra lookup ids and maintaining code to resolve them — whereas a graph just adds relationships followed during a query. It is the most natural operation in a graph. — Neo4j
Now the heart of the matter: what makes GraphRAG different from vector RAG.
slide 04 of 15
GraphRAG Theory: Local & Global
GraphRAG is a graph-based approach to question answering that “scales with both the generality of user questions and the quantity of source text.” It answers global sensemaking questions — the kind that naive RAG fails on. (Edge et al., arXiv:2404.16130)
Why naive RAG fails on global questions
Questions like “What are the main themes in this dataset?” are query-focused summarization (QFS) tasks, not explicit retrieval tasks. Retrieving a few similar chunks cannot answer them.
The two-tier architecture
GraphRAG has two query modes (Microsoft docs):
- Local search — retrieves granular entity/relationship evidence by traversing the graph.
- Global search — runs map-reduce over community summaries to synthesize corpus-wide answers.
Index construction
Building the index is a two-stage, LLM-driven pipeline:
Text
→ Entities & Relationships
→ Knowledge Graph
→ Graph Communities
→ Community Summaries
→ Community Answers
→ Global AnswerFirst the LLM derives an entity knowledge graph from the source documents. Then it pre-generates community summaries for all groups of closely related entities.
Community detection with Leiden
To find those communities, GraphRAG uses the Leiden algorithm (Traag et al., 2019). It exploits the graph’s inherent modularity to partition nodes into nested, modular communities of closely related entities — and GraphRAG recursively creates increasingly global summaries spanning this community hierarchy.
Answering with map-reduce
Global questions are answered by map-reduce: the map step generates partial answers independently and in parallel; the reduce step combines them into a final global answer.
The evaluation lens
The paper evaluates global sensemaking on:
- Comprehensiveness — how much detail, covering all aspects.
- Diversity — how varied the perspectives are.
- Empowerment — how well the answer helps the reader make informed judgments.
For global questions over ~1M-token datasets, GraphRAG gave substantial improvements over a conventional RAG baseline on both comprehensiveness and diversity. (Paper, Fig. 2)
Now we get hands-on. First: standing up a real graph database.
slide 05 of 15
Setup: Neo4j with Docker
Setting up Neo4j
Your options
- Neo4j Desktop (Linux, Mac, Windows)
- Docker container — the approach in this course
- A managed online instance
Docker Compose
Create a docker-compose.yaml:
services:
neo4j:
image: neo4j:5
ports:
- "7474:7474" # web UI
- "7687:7687" # Bolt protocol
environment:
NEO4J_AUTH: "neo4j/password123"The two ports
| Port | Purpose |
|---|---|
| 7474 | The Neo4j Browser web UI |
| 7687 | Bolt — the binary messaging protocol drivers use to talk to the DBMS |
Start it
docker-compose up -dThe first run sets the password. Neo4j is now on localhost — open the Browser UI at http://localhost:7474/browser.
What is Neo4j Browser?
A developer tool to execute Cypher queries and visualize the results — the default developer interface for both Enterprise and Community editions. It is where you will eyeball your graph.
Now the language of graphs: Cypher.
slide 06 of 15
Cypher Deep-Dive
Cypher: The Query Language of Graphs
Cypher is Neo4j’s declarative query language — it says what pattern you want, not how to find it.
The core clauses
| Clause | Role | Type |
|---|---|---|
| MATCH | Specifies patterns to find in the graph | Read |
| CREATE | Creates nodes and relationships | Write |
| MERGE | Ensures a pattern exists — matches or creates | Write |
| RETURN | Defines the result of the query | Read |
MERGE: the idempotent writer
In the course’s graph seeding we use MERGE everywhere. It matches an entity if it already exists, otherwise it creates it — so re-running the seed never duplicates data.
MERGE (f:Person {name: 'Florian', country: 'Austria'})
MERGE (c:YTChannel {name: 'NeuralNine'})
MERGE (os:OS {name: 'Linux'})
MERGE (l:Person {name: 'Linus Torvalds', country: 'Finland'})Relationships (edges)
MERGE (f)-[:OWNS]->(c)
MERGE (f)-[:USES]->(os)
MERGE (l)-[:CREATED]->(os)Each relationship is a first-class citizen, stored directly so it can be retrieved in one operation — that is what makes multi-hop queries fast.
A pattern-based read
MATCH (p:Person)-[:OWNS]->(:YTChannel {name: 'NeuralNine'})
RETURN p.nameWith a running database and a query language, let’s set up Python.
slide 07 of 15
Environment & Python Dependencies
Environment Setup
1. Get an OpenAI API key
Go to platform.openai.com/api-keys, create a secret key, and put it in a .env file:
OPENAI_API_KEY=your_key_here2. Install the Python packages
With pip:
pip install langchain-openai neo4j neo4j-graphrag python-dotenvOr with uv:
uv init
uv add langchain-openai neo4j neo4j-graphrag python-dotenvWhat each package is for
| Package | Role |
|---|---|
| langchain-openai | OpenAI chat models integrated with LangChain |
| neo4j | The official Neo4j Python driver (GraphDatabase) |
| neo4j-graphrag | GraphRAG retrievers, LLM wrappers, and schema helpers |
| python-dotenv | Loads OPENAI_API_KEY from .env |
Time to build the knowledge graph in code.
slide 08 of 15
Building the Knowledge Graph in Code
Connect to Neo4j
We connect over the Bolt protocol (port 7687) with the neo4j driver:
from neo4j import GraphDatabase
driver = GraphDatabase.driver(
'bolt://localhost:7687', auth=('neo4j', 'password123')
)Seed entities (nodes)
driver.execute_query("MERGE (f:Person {name: 'Florian', country: 'Austria'})")
driver.execute_query("MERGE (c:YTChannel {name: 'NeuralNine'})")
driver.execute_query("MERGE (os:OS {name: 'Linux'})")
driver.execute_query("MERGE (l:Person {name: 'Linus Torvalds', country: 'Finland'})")Seed relationships (edges)
driver.execute_query("MERGE (f)-[:OWNS]->(c)")
driver.execute_query("MERGE (f)-[:USES]->(os)")
driver.execute_query("MERGE (l)-[:CREATED]->(os)")Because these use MERGE they are safe to run many times — matches win over duplicates.
A tiny graph, a big idea
Even with four entities, we have already encoded real semantics: ownership, usage, and creation. Scale that to thousands of entities and you have a web of facts the agent can traverse.
Now the retriever that turns questions into graph queries.
slide 09 of 15
The Retriever: Text2CypherRetriever
The Text2CypherRetriever
The bridge between natural language and graph queries is Text2CypherRetriever from neo4j-graphrag (v1.18.0). It turns a natural-language query into Cypher via an LLM, then executes it against Neo4j.
The schema string
The retriever needs to know the graph’s structure so it writes valid Cypher. We provide it as text:
schema = """
Node labels:
- Person: name, country
- YTChannel: name
- OS: name
Relationships:
- Person -[:OWNS]-> YTChannel
- Person -[:USES]-> OS
- Person -[:CREATED]-> OS
"""This schema is injected into the prompt, instructing the LLM: do not use any properties or relationships not included in the schema. If omitted, the library auto-fetches it via get_schema(driver).
Create the retriever
from neo4j_graphrag.llm import OpenAILLM
from neo4j_graphrag.retrievers import Text2CypherRetriever
llm = OpenAILLM(model_name='gpt-4o-mini')
retriever = Text2CypherRetriever(
driver=driver, llm=llm, neo4j_schema=schema
)A built-in safety guard
This is a standout detail: before executing, Text2CypherRetriever runs EXPLAIN <query> and refuses to run anything whose query type is not read-only ("r"). Generated Cypher cannot mutate your graph. (Verified in the library source.)
Now let’s hand that retriever to an agent.
slide 10 of 15
Wiring the Agent: Tools & Streaming
Expose the retriever as a tool
We wrap the retriever in a @tool-decorated function so the LangChain agent can call it like any other tool:
from langchain.tools import tool
@tool
def query_kg(question: str) -> str:
"""Query the knowledge graph for information.
Pass entire user questions. Returns graph rows."""
results = retriever.search(question)
if not results.items:
return "No content"
return '\n'.join(item.content for item in results.items)Create the agent
from langchain.agents import create_agent
agent = create_agent(model='gpt-4o-mini', tools=[query_kg])Stream the reasoning
To see every step the agent takes, stream it:
for step in agent.stream(
{'messages': [('user', question)]},
stream_mode='values',
):
print(step['messages'][-1].content)The output reveals the chain of reasoning — the agent called the tool, got graph rows, and composed an answer from them.
Let’s watch that reasoning unfold on a genuinely multi-hop question.
slide 11 of 15
Multi-Hop Workshop: Trace the Path
A Multi-Hop Question, End to End
The question
Who runs NeuralNine, what operating system do they use, who created it, and where are they from?
The trace
Streaming the agent’s steps shows the exact path it walked through the graph:
Who runs NeuralNine? → Florian
What OS does Florian use? → Linux
Who created Linux? → Linus Torvalds
What country is Linus from? → FinlandEach hop is a MATCH across a relationship. A vector store could not answer this — no single chunk contains the whole chain.
Why this is explainable
This is GraphRAG’s superpower: the graph query replaces the “black box” of vector-only search with a detailed trace and the exact result set. The facts used can be linked back to the graph for auditing and re-ranking.
Now you try
- Add a second channel Florian owns.
- Ask “Who created the OS Florian uses?” and watch it hop twice.
- Add a relationship of your own and invent a multi-hop question.
Beyond the basics: how to combine graphs with vectors.
slide 12 of 15
Hybrid RAG: Vector + Graph Together
Why combine?
Graph RAG and vector RAG are complementary, not rivals. Neo4j’s text-based GraphRAG retrieval combines scored vector search with structural graph search: find relevant chunks by vector, extract their entities, then follow relationships multi-hop to reach farther, related information that pure vector search misses.
The vertical / horizontal idea
Documents are a vertical structure; entity extraction and clustering capture the horizontal re-occurrence of topics across documents. — Neo4j
Vectors live on graph nodes
In Neo4j’s GraphRAG for Python, entity names + descriptions, text chunks, and cluster summaries are each vector-embedded and stored in a vector index — and that index sits on the graph nodes, not in a separate store. Vector search here is approximate nearest neighbor, so it may not give exact results.
The canonical hybrid agent
The Neo4j + Milvus pattern routes each query to the vector store (semantic chunk similarity), the knowledge graph (entities/relationships), or both — with optional web search fallback and self-correction of hallucinations.
| Query type | Route |
|---|---|
| Simple similarity | Vector store |
| Entity / relationship | Knowledge graph |
| Multi-hop / global | Knowledge graph (local + global) |
| Ambiguous | Both, then merge |
Now let’s scale the whole thing to production.
slide 13 of 15
Scaling & Production
Scaling Up
The same code, a bigger graph
The code from this course works on far larger, richer datasets. Add more entities and relationships, switch to a more capable model, and the agent answers even more complex one-shot questions.
A genuinely hard example
On NeuralNine a lot of automation is taught. What is the capital of the country where the inventor of the programming language covered there was born?
The chain:
NeuralNine → teaches → Python
Python → invented by → Guido van Rossum
Guido van Rossum → born in → Netherlands
Netherlands → capital → AmsterdamWith a larger graph and a smarter model, the agent produces a single Cypher query that returns the answer directly: programming language Python, inventor Guido van Rossum, birth country Netherlands, capital Amsterdam.
Production considerations
- Schema refresh — the schema string auto-fetches/refreshes so the LLM writes valid Cypher as the graph grows.
- Entity resolution — de-duplicate entities via entity resolution, linking, and vector similarity.
- Read-only guardrails — the EXPLAIN check keeps generated Cypher read-only.
- Hybrid routing — a router sends each query to vector, graph, or both.
Where do real teams actually put this to work?
slide 14 of 15
Real-World Use Cases
Recommendation engines
Recommendations are naturally relational — “users who bought X also bought Y,” “people in your network like Z.” A graph stores those connections directly and walks them in one query.
Fraud & risk detection
Fraud rings are networks of accounts, devices, and transactions. Graph queries expose suspicious chains that isolated-record checks miss.
Biomedical & drug discovery
Genes, proteins, drugs, and diseases form a dense graph. Multi-hop queries reveal associations (drug → target → disease) that a flat search cannot.
Enterprise knowledge search
A company wiki is full of people, projects, and dependencies. GraphRAG gives grounded, explainable answers — and auditors love the traceable path.
When the graph is the answer
- Relationship-heavy, stable data
- Multi-hop and global questions
- Explainability and audit requirements
And the crucial counterpoint: when a graph is the wrong tool.
slide 15 of 15
Anti-Patterns & When Not to Use a Graph
Anti-Patterns & Pitfalls
When classic vector RAG wins
| Scenario | Best approach |
|---|---|
| Simple semantic / similarity search | Classic vector RAG |
| Isolated fact lookups | Classic vector RAG |
| Data changes very often | Classic vector RAG (graph reindex is costly) |
| Relationship-heavy queries & data | GraphRAG |
| Multi-hop reasoning over connected data | GraphRAG |
| Stable, highly interconnected data | GraphRAG |
Microsoft’s own caveat
Microsoft ships “Basic Search” as a GraphRAG query mode for queries best answered by baseline RAG (standard top-k vector search). Even the makers of GraphRAG route simple-similarity queries back to plain vectors.
Common mistakes
- Forcing a graph on flat data — if there are no meaningful relationships, you gain nothing.
- Ignoring the schema — a stale schema makes the LLM write Cypher against a graph that no longer matches.
- No entity resolution — duplicate entities fragment the graph and break multi-hop paths.
- Forgetting guardrails — rely on the read-only EXPLAIN check; never let generated Cypher write.
Vector RAG finds the most similar text; GraphRAG finds the connection. Choose the tool by the shape of the question.
You now have the full picture. Ready for the final assessment.
← → to move · g for contents