LightRAG vs GraphRAG is no longer a question of which one can update incrementally: GraphRAG 3.3.0 (8 October 2026) ships a graphrag update command and LightRAG 1.5.7 (2 September 2026) inserts and deletes documents in place, and both are MIT licensed. The real differences are three. GraphRAG needs a model that returns valid JSON for every community report (schema-checked) and for the map step of global search and DRIFT search, is most tested on OpenAI's gpt-4 series, and its own docs warn of frequent malformed JSON through Ollama or a LiteLLM proxy; LightRAG names Qwen3-30B-A3B-Instruct as a reasonable local minimum and does not require JSON at all. GraphRAG writes parquet tables and a LanceDB directory (or Azure Blob, Cosmos DB or AI Search), while LightRAG can keep KV, vectors, graph and document status in one PostgreSQL with no Apache AGE. GraphRAG's update adds new document titles only and appends new communities beside the old ones, so a withdrawn document means a full re-index; LightRAG deletes a document and rebuilds the affected entities from its extraction cache instead of re-extracting. Pick GraphRAG for thematic questions over a mostly static corpus, LightRAG for entity lookups over content that gets revised or withdrawn. This week, run both extraction pipelines on 20 to 50 representative documents with the exact local model you will deploy.
GraphRAG 3.3.0 shipped on 8 October 2026 with a graphrag update command and two incremental indexing methods. That matters for anyone comparing LightRAG vs GraphRAG, because the 2024 LightRAG paper set the terms of this comparison with two claims: that GraphRAG has to rebuild its whole graph on every change, and that LightRAG retrieves with thousands of times fewer tokens. The first is out of date. The second is the paper authors' own estimate against an older GraphRAG, and it does not describe your corpus.
The question an engineer at a bank, insurer or pharma company actually has is narrower: which of these two can I run on the models and databases I already operate, and what happens to the graph when a document is revised or withdrawn? Both projects are MIT licensed and actively shipping (LightRAG at roughly 40,000 GitHub stars, Microsoft's GraphRAG at roughly 36,000), so licensing settles nothing. The choice turns on whether your local model can return valid JSON, whether the graph has to live in a database you already govern, and whether documents ever get revised or withdrawn.
Everything below is pinned to GraphRAG 3.3.0 and LightRAG 1.5.7 and comes from the two projects' documentation and source at those tags, plus the LightRAG paper. Nothing was benchmarked for this post. If you are still deciding whether you need graph RAG at all, start with our comparison of cache-augmented generation and GraphRAG, and for how GraphRAG's entity extraction and community structure work, see our GraphRAG implementation retrospective. This post does not re-explain either.
LightRAG vs GraphRAG at a glance
The rest of this post is the reasoning behind each row that changes a decision.
| GraphRAG 3.3.0 | LightRAG 1.5.7 | |
|---|---|---|
| Licence | MIT | MIT |
| Latest release | 8 Oct 2026 | 2 Sep 2026 |
| Indexing LLM work | Entity and relationship extraction, summarisation, community reports | Per-chunk entity and relationship extraction, plus summaries for heavily merged entities |
| Steps that need valid JSON | Community reports (standard and fast), global search, DRIFT search | None required; JSON extraction is optional |
| Stated local model guidance | None; most tested on OpenAI gpt-4 series | Qwen3-30B-A3B-Instruct as a reasonable minimum, non-thinking |
| Query modes | local, global, drift, basic | local, global, hybrid, naive, mix (default) |
| Default storage | Parquet files plus LanceDB on local disk | JSON, NanoVectorDB and NetworkX files, not for production |
| Production storage options | File, Azure Blob or Cosmos DB; LanceDB, Azure AI Search or Cosmos DB for vectors | PostgreSQL, MongoDB or OpenSearch as one backend; Neo4j or Memgraph plus Milvus or Qdrant |
| Incremental add | graphrag update, new titles only | Insert by content or file hash |
| Delete or revise a document | No; full re-index | adelete_by_doc_id, rebuild from extraction cache |
| Offline prerequisites | NLTK corpora and spaCy models for fast mode | Optional pip packages and tiktoken files, or the Docker image |
Indexing cost: where the LLM calls go
GraphRAG's own documentation estimates graph extraction at roughly 75% of indexing cost, and its README opens with a warning that indexing "can be an expensive operation" and tells you to start small. The standard pipeline calls the model to extract entities and relationships from every chunk, to summarise entity and relationship descriptions, and then to write a report for every community. Budget for that last step separately: it runs in both indexing methods.
FastGraphRAG (graphrag index --method fast) replaces LLM extraction with noun-phrase extraction (NLTK plus regex by default, spaCy options built in), derives relationships from co-occurrence within a text unit, skips summarisation, and still uses the model for community reports. The docs say they generally configure it with much smaller chunks, 50 to 100 tokens. The documentation is candid about the trade: the graph "tends to be quite a bit noisier" and is less useful outside GraphRAG, but if your use case is mainly summary questions through global search, fast mode gives "high quality summarization with much lower language model cost". The default NLTK extractor is described as primarily suitable for English, which matters for a multilingual document estate. LazyGraphRAG, Microsoft Research's deferred-extraction approach, is a different thing: it is not an indexing or search method in the open-source package at 3.3.0. Our post on LazyGraphRAG and cheaper knowledge-graph indexing covers the idea.
LightRAG spends most of its indexing budget on per-chunk entity and relationship extraction with a non-thinking model, including one gleaning pass per chunk by default, and calls the model again only to summarise an entity or relationship once its descriptions pile up (8 by default, FORCE_LLM_SUMMARY_ON_MERGE). There is no community-report stage. The extraction results are cached by default, and that cache is what later makes deletion cheap.
Neither project publishes per-corpus call counts or costs for current releases, so this post does not invent any. The number that circulates is from the LightRAG paper (arXiv 2410.05779). On one Legal dataset of 94 documents and about 5.08 million tokens, the authors estimate that 2024 GraphRAG global search reads 610 level-2 community reports at about 1,000 tokens each, around 610,000 tokens over hundreds of API calls, against fewer than 100 tokens in one call for LightRAG's keyword generation and retrieval. GPT-4o-mini was the default model and also the judge, with quality scored as LLM-as-judge pairwise win rates. That is the authors' own analytic estimate for one query mode on one corpus, not an independent benchmark. The paper's update-cost argument also assumes GraphRAG regenerates community reports on update, which 3.x does not do for existing communities (see the updates section below).
The cost decision is therefore structural, not numeric. GraphRAG charges you up front for community reports so that global questions are answerable later. LightRAG charges less up front and answers whole-corpus questions less directly.
Running each on a local model
This is the section that decides the comparison for an on-prem team, because it is where the two projects differ most.
GraphRAG: JSON is the hard requirement
GraphRAG calls every model through LiteLLM, and its model documentation says a model "must support returning structured outputs adhering to a JSON schema". It has been most thoroughly tested on the gpt-4 series: gpt-4, gpt-4-turbo, gpt-4o and gpt-4o-mini. The quality evaluation in GraphRAG's own paper used gpt-4-turbo (arXiv 2404.16130). On local models routed through Ollama or a LiteLLM proxy, the docs say the setup "seems to work reasonably well, but we frequently see issues with malformed responses (especially JSON)". Precision matters here, because the dependency is not in extraction. Graph extraction at 3.3.0 uses a delimited tuple prompt, not JSON. Community report generation passes a JSON schema as the response format and parses the reply against it, the map step of global search requests a JSON object, and DRIFT search parses JSON from the model too. A local model that cannot return valid JSON therefore leaves communities without reports in both standard and fast indexing, and it degrades two of the four query modes: global search logs a parsing warning and skips each map batch it cannot parse. If you serve the model through vLLM, our post on why self-hosted structured output fails silently covers the server side of that problem. GraphRAG names no local model that satisfies its JSON requirement, so you have to prove one yourself. Custom model classes can be registered, but the docs say that is not supported with the CLI and requires using GraphRAG as a library.
LightRAG: a stated minimum model and no JSON dependency
LightRAG's README splits the work into roles. The extraction model should be non-thinking (reasoning disabled), and for local deployment "Qwen3-30B-A3B-Instruct is a reasonable minimum". The keyword model must be non-thinking. The query model should be stronger than the extraction model. Role-specific variables such as EXTRACT_LLM_MODEL and EXTRACT_LLM_BINDING let you point extraction at a different server from querying. The LLM_BINDING options include openai, which covers any OpenAI-compatible server, and ollama. ENTITY_EXTRACTION_USE_JSON defaults to false in code, so extraction uses a delimited text format unless you turn JSON on. The shipped example environment file turns it on, noting that JSON costs latency but improves reliability. Either way, nothing in the pipeline depends on schema compliance the way GraphRAG's community reports do. The README is equally specific about where local extraction breaks. It names 3 causes of extraction timeouts: a model below roughly 50 tokens per second, chunks that produce too many entities, and output loops, with locally deployed Qwen models called out for the last. The effective timeout is twice the configured value (EXTRACT_LLM_TIMEOUT=300 allows up to 600 seconds), and the README's suggested sizing rule is to keep max output tokens below the timeout multiplied by tokens per second, for example 9,000 under 240 seconds at 50 tokens per second.
Where the graph lives: parquet files vs your Postgres
GraphRAG 3.3.0 writes its output tables as parquet files on disk by default (CSV and Cosmos DB are the other table types), its vectors to a LanceDB directory under the output folder (Azure AI Search and Cosmos DB are the alternatives), and its storage can be file, memory, Azure Blob or Cosmos DB. No PostgreSQL, Neo4j or other graph database backend ships in the package. If your estate is Azure Blob and Cosmos DB, that fits. If it is not, the graph is a directory of files you have to back up, encrypt and access-control yourself.
LightRAG's defaults (JsonKVStorage, JsonDocStatusStorage, NetworkXStorage, NanoVectorDBStorage) are, in the 1.5.7 README's words, "intended only for development and debugging, and are not suitable for production." The single-backend production options are PostgreSQL, MongoDB or OpenSearch, and the specialised options are Neo4j or Memgraph for the graph and Milvus or Qdrant for vectors.
Postgres is the option most regulated teams already run under existing controls, and two details decide whether it works on yours. First, the graph does not need Apache AGE: PGTableGraphStorage runs on plain PostgreSQL 14 or later tables with no extensions. PGGraphStorage uses AGE and refuses to start on AGE 1.8.0 or newer, so use it only if you need AGE. Second, vectors do need pgvector: PGVectorStorage runs CREATE EXTENSION IF NOT EXISTS vector. On a database where your DBA has not made the extension available, that statement fails, LightRAG only logs a warning, and the vector tables, which use the vector column type, cannot be created.
The configuration below puts all four LightRAG stores in one Postgres and points the models at local OpenAI-compatible servers. Hosts are illustrative, and LLM_MODEL must match the model id your server actually exposes.
# lightrag-server environment file (v1.5.7) # Storage: everything in the Postgres you already run LIGHTRAG_KV_STORAGE=PGKVStorage LIGHTRAG_DOC_STATUS_STORAGE=PGDocStatusStorage LIGHTRAG_VECTOR_STORAGE=PGVectorStorage # needs the pgvector extension LIGHTRAG_GRAPH_STORAGE=PGTableGraphStorage # plain tables, no Apache AGE POSTGRES_HOST=pg.internal POSTGRES_PORT=5432 POSTGRES_USER=lightrag POSTGRES_PASSWORD='change-me' POSTGRES_DATABASE=rag # Default LLM for all roles: local, non-thinking LLM_BINDING=openai LLM_BINDING_HOST=http://llm.internal:8000/v1 LLM_MODEL=Qwen3-30B-A3B-Instruct ENTITY_EXTRACTION_USE_JSON=true # Embeddings: pick once, before the first insert EMBEDDING_BINDING=openai EMBEDDING_BINDING_HOST=http://embed.internal:8000/v1 EMBEDDING_MODEL=BAAI/bge-m3 EMBEDDING_DIM=1024 # Local reranker served by vLLM (cohere-compatible binding) RERANK_BINDING=cohere RERANK_MODEL=BAAI/bge-reranker-v2-m3 RERANK_BINDING_HOST=http://rerank.internal:8000/rerank # Offline: pre-seeded tiktoken files TIKTOKEN_CACHE_DIR=/opt/lightrag/tiktoken
EMBEDDING_DIM=1024 is a property of bge-m3's dense output, not of LightRAG. Get it wrong and the first table creation bakes in the wrong dimension, which is the next section's problem. The reranker is off by default (RERANK_BINDING=null); the example file says a reranker served by vLLM should use the cohere binding.
Incremental updates, deletes and embedding changes
Both tools update incrementally. The difference is what an update can express.
# settings.yaml (GraphRAG 3.3.0) output_storage: type: file base_dir: "output" update_output_storage: # per-run backup of the old index + the delta type: file base_dir: "update_output" cache: type: json # re-runs reuse cached LLM responses # then, for new documents only (matched by title): # graphrag update --root ./project # graphrag update --root ./project --method fast # FastGraphRAG delta
What graphrag update does
graphrag update takes the same --root and --method options as indexing, and uses the LLM cache by default. At 3.3.0 the pipeline creates a timestamped folder under update_output_storage, copies the current index into a previous child as a backup, indexes only the new documents into a delta child, merges previous and delta, and writes the merged tables back to the main output_storage. The live index is still output_storage; update_output_storage holds per-run backups and deltas, which is what the docs mean by "a secondary storage location for running incremental indexing, to preserve your original outputs." Three source-level details shape how you operate it: The practical rule: use graphrag update for additions between scheduled full re-indexes, and schedule the full re-index, because the community structure drifts from what a clean run would produce. The configuration below keeps the backups on separate storage. Remember that it only ever adds new titles; it neither deletes nor revises.
- 1. New means a new title. A document is treated as new only if its title is not already in the documents table. The code also computes which titles disappeared from the input, but nothing in the package reads that list. From the source, the consequence is that a revised document that keeps its title is not reprocessed, and a removed document is not removed from the graph.
- 2. Communities are appended, not re-clustered. Delta communities get ids offset past the old maximum and are concatenated onto the existing ones, and community reports are merged the same way. Existing communities are not re-clustered with the new entities, and their reports are not regenerated. Entities are resolved across old and delta by title.
- 3. There is no delete. Removing a document's contribution means a full re-index.
What LightRAG does
LightRAG identifies documents by an MD5 id generated from content or file path, drops duplicates and skips documents it has already processed, and inserts new ones into the live stores. Deletion is a first-class operation: adelete_by_doc_id(doc_id, delete_llm_cache=False) removes a document, and where entities or relationships were shared with other documents, they are rebuilt from the cached extractions of the remaining documents rather than by re-running extraction. Deletion is serialised against the pipeline: while an insert job runs, a delete is rejected with status not_allowed, not queued, so your withdrawal job has to retry. A revised document is handled as delete then re-insert, which LightRAG supports and GraphRAG's update does not.
Embedding changes
LightRAG's README is blunt: the embedding model must be chosen before indexing, used again at query time, and changing it means re-embedding all chunks, entities and relationships. In PostgreSQL the vector dimension is fixed when the tables are first created. Our explainer on embedding dimensions in vector search covers why that choice is sticky. The same README paragraph says LightRAG has no re-embedding tool, but the lightrag-rebuild-vdb tool documentation contradicts it and is the one to trust: it rebuilds the entity, relationship and chunk vector stores from the graph and the stored text chunks, and supports rebuilding everything after an embedding model or dimension change. Stop the LightRAG server first. You pay the embedding cost again, but no LLM extraction is re-run. For generic vector-index update patterns that apply to either tool, see our guide to updating RAG knowledge without rebuilding.
Query modes mapped to question types
GraphRAG describes global search as resource-intensive, and it is the mode that needs a JSON-reliable model on every map call. LightRAG says mix takes slightly longer than naive and the other modes are roughly comparable in latency. Adding its recommended reranker, BAAI/bge-reranker-v2-m3 deployed locally, typically adds 1 to 2 seconds per query, and the rerank model can be changed at any time without touching the index. Whether that delay pays for itself is the question our guide to when reranking is worth it answers.
The table decides the comparison by the questions your users ask. If they mostly ask whole-corpus thematic questions, GraphRAG's community reports are the asset you are paying for. If they mostly ask about named entities and the links between them, LightRAG answers those without the community-report bill.
| Question shape | GraphRAG mode | LightRAG mode |
|---|---|---|
| Themes across the whole corpus ("what risks recur across all supplier contracts?") | global: map-reduce over all community reports, resource-intensive | global: relationship chains across broad themes, no community reports |
| A specific entity and its context ("what do we know about counterparty X?") | local: graph data plus raw text chunks | local: candidate entities and their attributes |
| How entities relate ("which suppliers share a parent company?") | local or drift | global (relationship chains) or hybrid (local plus global) |
| Entity question that needs broader context | drift: local search plus community information and follow-up questions | mix (default): local, global and naive combined |
| Plain passage lookup | basic: rudimentary vector RAG | naive: vector similarity over text chunks, no graph |
On-prem and air-gapped: what changes the answer
For a team whose weights and data stay inside the perimeter, the general comparison above narrows to four checks.
Model fit is the binding constraint. With a hosted gpt-4-class model, GraphRAG's JSON dependency is invisible. With a mid-size model on your own GPUs it is the first thing to test, because a failure lands at community reports, which every GraphRAG index needs. LightRAG publishes a local minimum and does not need JSON. If the largest model you can serve at a usable speed is around the 30B-A3B class, LightRAG is the tool with documented expectations for it.
The graph should live under controls you already evidence. LightRAG on one PostgreSQL inherits that database's backups, encryption at rest, access control and audit logging. GraphRAG's parquet tables and LanceDB directory are a filesystem path that needs its own answer to each of those, unless you are an Azure Blob or Cosmos DB shop.
Offline installs need pre-staging for both. FastGraphRAG downloads NLTK corpora (brown, treebank, averaged_perceptron_tagger_eng, punkt, punkt_tab) and any missing spaCy model on first use. LightRAG installs optional storage clients and LLM SDKs at runtime and downloads tiktoken files from OpenAI's CDN, all of which fail offline. Its fix is pip install lightrag-hku[offline] on a connected machine, TIKTOKEN_CACHE_DIR, and lightrag-download-cache, or the LightRAG Docker image, which is pre-configured for offline use. The model-server half of the air gap is covered in our vLLM air-gapped deployment guide.
Update semantics are a compliance question. When a policy is superseded, a contract is terminated or a data subject's records must be removed, LightRAG can delete the document and regenerate the affected entities from cache. GraphRAG's update cannot express removal, so honouring a withdrawal means a full re-index, and until it runs the old content is still answerable. For regulated content, that gap can decide the choice on its own.
Decision table and what to do this week
For the wider map of retrieval architectures, the RAG systems pillar puts graph RAG in context.
This week, take 20 to 50 representative documents and run both extraction pipelines on them with the exact local model you intend to deploy, not a hosted stand-in. For GraphRAG, check the indexing log for "error generating community report" and the query log for "Error parsing search response json" warnings, since a global query still returns an answer when map batches are skipped. For LightRAG, count how many chunks hit the extraction timeout and whether any show output loops. Then delete one document from the LightRAG index and confirm its entities are gone. Those three results will decide this comparison for your stack more reliably than any published token count.
| Situation | Choose | Why |
|---|---|---|
| Thematic, whole-corpus questions over a mostly static corpus | GraphRAG standard, or --method fast if questions are mainly global | Community reports are what make global answers possible |
| Entity and relationship lookups over content that is revised or withdrawn | LightRAG on Postgres | Hash-based insert plus delete with cached rebuild |
| Only mid-size local models available, JSON compliance unproven | LightRAG | No JSON-schema dependency; documented 30B-A3B minimum |
| Estate built on Azure Blob, Cosmos DB or AI Search, gpt-4-class model available | GraphRAG | Storage and model match what it was built and tested on |
| Graph must live in an existing governed PostgreSQL | LightRAG | PGTableGraphStorage plus pgvector, no AGE |
| Long, structured single documents (filings, manuals) | Neither; see PageIndex and vectorless RAG | Document structure matters more than a cross-document graph |
FAQ
Quick answers to the questions this post tends to raise.



