kmail.at
← learning

technologies · difficulty ◆◆◆

Graphiti — temporal knowledge graphs for AI agents

A knowledge graph that remembers what changed — and when.

Most knowledge graphs are a snapshot: they tell you what is true now, and forget everything that came before. Graphiti is built the other way — every fact carries a validity window, so an AI agent can ask not just "what is true?" but "what was true in March, and when did it change?"

2026-08-14 · 16 min read

$ graphiti-core

What it is

Graphiti is a Python framework for building and querying temporal context graphs for AI agents. A context graph is a graph of entities, relationships, and facts where every fact has a validity window — when it became true, and when (if ever) it was superseded. Unlike a static knowledge graph, Graphiti tracks how facts change over time, keeps provenance back to the raw source data, and supports both prescribed and learned ontology. It is the open-source engine at the core of Zep’s agent-memory infrastructure.

Why it matters

Traditional retrieval-augmented generation (RAG) treats knowledge as flat chunks of text, and traditional knowledge graphs treat it as a static snapshot. Both break when the world changes: a fact that was true yesterday is stale today, and a flat chunk cannot tell you when it stopped being true. Graphiti solves this by continuously integrating new data into a graph, invalidating old facts instead of deleting them, and preserving the full temporal history. An agent can answer "what is true now" and "what was true at any point in time" — which is exactly what a memory system needs.

How it fits your stack

Graphiti ingests both unstructured text and structured JSON as episodes. Each episode is processed to extract entities (nodes) and relationships (edges), which are merged into a unified graph. You query that graph with hybrid search that fuses semantic embeddings, BM25 keyword retrieval, and graph traversal. The 0.29.x line added a big workflow win: the falkordblite extra ships an embedded FalkorDB that runs in-process with no Docker at all (Python 3.12+). There is also a first-class MCP server so Claude, Cursor, and other MCP clients can use context-graph memory directly. It runs against a graph database you bring — Neo4j, FalkorDB, Amazon Neptune, or (deprecated) Kuzu — and defaults to OpenAI for LLM inference and embedding, with Anthropic, Gemini, Groq, and OpenAI-compatible local models supported.

Example

$ pip install "graphiti-core[falkordblite]"
  Collecting graphiti-core
  Downloading graphiti_core-0.29.3-py3-none-any.whl
  Collecting falkordblite
  Installing collected packages: graphiti-core, falkordblite
  Successfully installed graphiti-core-0.29.3

Install the core library with the embedded FalkorDB Lite backend — everything runs in-process, so you can skip a separate graph database entirely.

$ docker run -p 6379:6379 -it --rm falkordb/falkordb:latest
  Unable to find image 'falkordb/falkordb:latest' locally
  latest: Pulling from falkordb/falkordb
  ...
  FalkorDB server started on port 6379

Or, if you prefer a dedicated graph server (and no Python 3.12 constraint), FalkorDB in Docker is still the fastest containerized option — no Neo4j Desktop install needed.

$ python quickstart_neo4j.py
  Added episode: Freakonomics Radio 0 (text)
  Added episode: Freakonomics Radio 1 (text)
  Added episode: Freakonomics Radio 2 (json)
  Added episode: Freakonomics Radio 3 (json)
  
  Searching for: 'Who was the California Attorney General?'
  Search Results:
  UUID: 3f2a...
  Fact: Kamala Harris was the Attorney General of California
  Valid from: 2011-01-03
  Valid until: 2017-01-03

The quickstart ingests text and JSON episodes, then runs a hybrid search. Notice the result carries a validity window — the fact knows when it was true. The repo also ships quickstart_falkordb.py and quickstart_neptune.py for the other backends.

$ graphiti.search(query, center_node_uuid=node_uuid)
  Reranking search results based on graph distance:
  Using center node UUID: 3f2a...
  Reranked Search Results:
  UUID: 9c1b...
  Fact: Gavin Newsom is the Governor of California
  Valid from: 2019-01-07

Passing a center node reranks results by graph distance — facts closer to the focal node rank higher, which is how you center search on a person or topic.

Common flags

add_episode()
ingest a text or JSON episode into the graph
search()
hybrid search over relationships (semantic + BM25 + graph)
center_node_uuid
rerank results by graph distance to a focal node
_search()
lower-level search with a config recipe (e.g. NODE_HYBRID_SEARCH_RRF)
falkordblite
pip extra that runs an embedded FalkorDB in-process (Python 3.12+)
MCP server
Model Context Protocol server exposing memory tools over HTTP at /mcp/
SEMAPHORE_LIMIT
env var controlling ingestion concurrency (default 10)
graph_driver
pass a custom Neo4j/FalkorDB driver for a custom database

History

Origin

Graphiti was created by Zep AI, the company behind Zep’s context infrastructure for AI agents. It was open-sourced in August 2024 as the temporal context graph engine at the core of Zep’s platform. The team published the underlying architecture in a paper, "Zep: A Temporal Knowledge Graph Architecture for Agent Memory" (arXiv 2501.13956), and demonstrated that Zep is the State of the Art in Agent Memory. Graphiti is the open-source core, while Zep is the managed, production-grade platform built on top of it.

The temporal insight

The core idea borrows from how databases handle time: instead of overwriting a fact when it changes, you mark the old one invalid and add the new one, each with a validity window. This is bi-temporal tracking — the graph knows both when a fact was true in the world and when the system learned it. That single design decision is what lets an agent answer historical questions and reconcile contradictions automatically, without an LLM having to judge them.

Growth & adoption

Graphiti has grown rapidly since its release, reaching roughly 30,000 GitHub stars and becoming a reference point for agent memory and temporal knowledge graphs. It ships an MCP server so AI assistants like Claude and Cursor can use context-graph memory, a FastAPI REST service, and a rich set of search recipes. The 0.29.x releases added the embedded FalkorDB Lite backend, richer search recipes (MMR, cross-encoder reranking, node-distance, community search), and richer entity types for structured knowledge extraction.

Fun facts

Pros & cons

pros

  • + Temporal facts with validity windows — query what was true at any point in time
  • + Full provenance: every fact traces back to its source episode
  • + Embedded FalkorDB Lite backend means you can run it with no Docker at all
  • + Hybrid retrieval (semantic + BM25 + graph) for high-precision, sub-second queries
  • + MCP server gives Claude, Cursor, and other assistants direct memory access
  • + Incremental ingestion — the graph evolves in real time without batch recomputation

cons

  • − Best results need an LLM with Structured Output support; small models can fail ingestion
  • − Self-hosted only — you build and operate the surrounding system
  • − Embedded FalkorDB Lite requires Python 3.12+
  • − Younger and less battle-tested than traditional RAG pipelines

Takeaways

  1. 1Try it fastest: `pip install "graphiti-core[falkordblite]"` and run an in-process embedded graph (Python 3.12+)
  2. 2Or use a dedicated backend: FalkorDB in Docker, or Neo4j for the classic quickstart
  3. 3Ingest both text and JSON as episodes — Graphiti extracts entities and relationships automatically
  4. 4Query with `graphiti.search()` for hybrid retrieval that fuses semantic, keyword, and graph signals
  5. 5Pass a `center_node_uuid` to rerank results by graph distance for focused, context-aware search
  6. 6Wire it into an assistant with the MCP server (tools like add_memory and search_memory_facts)

Related commands

← all learning