Graphmemdocs GitHub ↗

Retrieval evaluation results

Concrete numbers for gmem recall on three multi-hop QA benchmarks. For the interpretation, see conclusions.md .

Setup

  • Scripts: scripts/eval_retrieval.py drives the gmem mcp server over stdio; graphs come from scripts/extract_eval_graphs.py (spaCy en_core_web_sm, no LLM). Datasets come from scripts/download_eval_datasets.py.
  • Datasets: HotpotQA (distractor), 2WikiMultihopQA and MuSiQue, validation split, first 100 questions each.
  • Corpus: shared (default) — every paragraph of the selected questions, deduplicated, in one scope; each question ranks the whole corpus, as HippoRAG evaluates. (per-question keeps each question’s own paragraphs, where recall@10 is ~1 and nothing is discriminated.)
  • Baseline model: sentence-transformers/msmarco-distilbert-cos-v5.
  • Binary: release gmem built with --features cuda, GRAPHMEM_EMBEDDING_BACKEND=auto → CUDA (RTX 4060), batch 16. The 50-question run reproduces the earlier CPU batch-1 numbers exactly, so the backend changes speed, not quality.
  • Metrics: mean over questions. recall@k = fraction of supporting paragraphs in the top k; MRR = 1 / rank of the first supporting paragraph.

Graphs:

  • none: no entities.
  • mentions: one entity per paragraph title, plus an edge when a paragraph’s text mentions another title from the same question.
  • spacy: spaCy entities and subject-verb-object triples, corpus-filtered (--spacy-max-df 0.02 --spacy-min-df 2).
  • oracle: the dataset’s own gold triples. These cover only the answer path, so this leaks the answer and is only a ceiling.

Modes:

  • embeddings: recall with use_embeddings=true (semantic seeds + Personalized PageRank).
  • fts-raw: use_embeddings=false with the raw question.
  • fts-or: use_embeddings=false with the question’s words OR-joined.

MS MARCO DistilBERT: 100 questions, shared corpus

datasetgraphmoderecall@2recall@5recall@10MRR
HotpotQAnoneembeddings0.5800.7250.8300.889
HotpotQAnonefts-raw0.5400.7400.8950.872
HotpotQAnonefts-or0.5500.7450.8950.885
HotpotQAmentionsembeddings0.6450.8250.9400.896
HotpotQAspacyembeddings0.6200.8000.9200.887
2Wikinoneembeddings0.6150.6980.7280.953
2Wikinonefts-raw0.6050.6900.7550.950
2Wikinonefts-or0.6050.6900.7550.950
2Wikimentionsembeddings0.7230.9150.9680.962
2Wikispacyembeddings0.7050.8330.9070.980
2Wikioracleembeddings0.8280.9630.9650.985
MuSiQuenoneembeddings0.4900.5800.6550.858
MuSiQuenonefts-raw0.4650.5200.5650.812
MuSiQuenonefts-or0.4650.5200.5650.812
MuSiQuementionsembeddings0.5600.6850.7400.888
MuSiQuespacyembeddings0.6050.7350.8550.891
MuSiQueoracleembeddings0.7650.8550.8850.930

IBM Granite Embedding: 100 questions, shared corpus

Run with GRAPHMEM_EMBEDDING_MODEL=ibm-granite/granite-embedding-97m-multilingual-r2, GRAPHMEM_EMBEDDING_BACKEND=auto, and ~/.cargo/bin/gmem (CUDA-enabled) on an NVIDIA GeForce RTX 4060 Laptop GPU. The harness’s standard matrix (all graphs and modes above, 100 questions per dataset) completed in about 90 seconds.

datasetgraphmoderecall@2recall@5recall@10MRR
HotpotQAnoneembeddings0.6450.8250.9350.911
HotpotQAnonefts-raw0.5400.7400.8950.872
HotpotQAnonefts-or0.5500.7450.8950.885
HotpotQAmentionsembeddings0.6650.8700.9650.910
HotpotQAspacyembeddings0.6350.8400.9350.904
2Wikinoneembeddings0.6670.7450.7770.985
2Wikinonefts-raw0.6050.6900.7550.950
2Wikinonefts-or0.6050.6900.7550.950
2Wikimentionsembeddings0.7050.8250.9100.990
2Wikispacyembeddings0.6850.7620.8070.990
2Wikioracleembeddings0.8450.9680.9680.995
MuSiQuenoneembeddings0.5000.6250.6900.898
MuSiQuenonefts-raw0.4650.5200.5650.812
MuSiQuenonefts-or0.4650.5200.5650.812
MuSiQuementionsembeddings0.5300.6800.7450.922
MuSiQuespacyembeddings0.5300.6650.7450.914
MuSiQueoracleembeddings0.7850.8950.9250.977

The lexical rows are identical to the MS MARCO run. Among non-oracle graph settings, mentions leads HotpotQA and 2Wiki recall; on MuSiQue, mentions and spacy tie at recall@10, with mentions ahead at recall@5 and MRR. Oracle results leak gold reasoning paths and are only a ceiling.

all-MiniLM-L6-v2: 100 questions, shared corpus

Run at commit 0db8dcf with GRAPHMEM_EMBEDDING_MODEL=sentence-transformers/all-MiniLM-L6-v2, the same CUDA release configuration, RTX 4060, batch 16, corpus, graph extraction, and first 100 validation questions as the DistilBERT baseline above. The complete sweep (11 reembeds) took 104 s.

datasetgraphmoderecall@2recall@5recall@10MRR
HotpotQAnoneembeddings0.1550.2150.2500.300
HotpotQAnonefts-raw0.5400.7400.8950.872
HotpotQAnonefts-or0.5500.7450.8950.885
HotpotQAmentionsembeddings0.3450.4850.6200.571
HotpotQAspacyembeddings0.3800.5250.6950.639
2Wikinoneembeddings0.0330.0480.0480.065
2Wikinonefts-raw0.6050.6900.7550.950
2Wikinonefts-or0.6050.6900.7550.950
2Wikimentionsembeddings0.4530.5800.6600.806
2Wikispacyembeddings0.5570.6570.6880.924
2Wikioracleembeddings0.6820.8630.8850.882
MuSiQuenoneembeddings0.0650.0900.1000.129
MuSiQuenonefts-raw0.4650.5200.5650.812
MuSiQuenonefts-or0.4650.5200.5650.812
MuSiQuementionsembeddings0.2950.3550.3950.510
MuSiQuespacyembeddings0.3050.3800.4250.558
MuSiQueoracleembeddings0.4600.5600.7050.733

MiniLM is faster for this sweep but its embedding recall is lower than the MS MARCO DistilBERT baseline in every non-oracle setting. The lexical rows are unchanged because they do not use embeddings. As with the baseline, graph context substantially improves MiniLM over its none configuration.

all-MiniLM-L12-v2: 100 questions, shared corpus

Run at commit 0db8dcf with GRAPHMEM_EMBEDDING_MODEL=sentence-transformers/all-MiniLM-L12-v2, the same CUDA release configuration, RTX 4060, batch 16, corpus, graph extraction, and first 100 validation questions as the other runs. The complete sweep (11 reembeds) took 217 s.

datasetgraphmoderecall@2recall@5recall@10MRR
HotpotQAnoneembeddings0.0950.1200.1250.186
HotpotQAnonefts-raw0.5400.7400.8950.872
HotpotQAnonefts-or0.5500.7450.8950.885
HotpotQAmentionsembeddings0.2800.3950.4700.482
HotpotQAspacyembeddings0.3250.4700.5700.546
2Wikinoneembeddings0.0180.0280.0280.037
2Wikinonefts-raw0.6050.6900.7550.950
2Wikinonefts-or0.6050.6900.7550.950
2Wikimentionsembeddings0.4500.5420.5900.815
2Wikispacyembeddings0.5930.6570.6670.952
2Wikioracleembeddings0.7250.8530.8600.914
MuSiQuenoneembeddings0.0100.0200.0250.021
MuSiQuenonefts-raw0.4650.5200.5650.812
MuSiQuenonefts-or0.4650.5200.5650.812
MuSiQuementionsembeddings0.2300.2800.3050.432
MuSiQuespacyembeddings0.2800.3150.3450.539
MuSiQueoracleembeddings0.4000.4200.4550.759

L12 is about twice as slow as L6 and lower on HotpotQA and MuSiQue. Its 2Wiki spaCy recall@2 improves from 0.557 to 0.593, but recall@5 falls to 0.657 and recall@10 to 0.667. L6 remains the better general-purpose MiniLM choice for this retrieval workload.

msmarco-MiniLM-L6-cos-v5: 100 questions, shared corpus

Run at commit 0db8dcf with GRAPHMEM_EMBEDDING_MODEL=sentence-transformers/msmarco-MiniLM-L6-cos-v5, the same CUDA release configuration, RTX 4060, batch 16, corpus, graph extraction, and first 100 validation questions as the other runs. The complete sweep (11 reembeds) took 112 s.

datasetgraphmoderecall@2recall@5recall@10MRR
HotpotQAnoneembeddings0.5250.6850.8050.850
HotpotQAnonefts-raw0.5400.7400.8950.872
HotpotQAnonefts-or0.5500.7450.8950.885
HotpotQAmentionsembeddings0.5700.7950.9100.850
HotpotQAspacyembeddings0.5700.7550.8800.864
2Wikinoneembeddings0.6030.6800.7250.970
2Wikinonefts-raw0.6050.6900.7550.950
2Wikinonefts-or0.6050.6900.7550.950
2Wikimentionsembeddings0.7300.9100.9730.985
2Wikispacyembeddings0.7100.8300.9020.995
2Wikioracleembeddings0.8280.9580.9680.995
MuSiQuenoneembeddings0.4700.5350.6200.874
MuSiQuenonefts-raw0.4650.5200.5650.812
MuSiQuenonefts-or0.4650.5200.5650.812
MuSiQuementionsembeddings0.5550.6500.6950.886
MuSiQuespacyembeddings0.6050.7500.8500.891
MuSiQueoracleembeddings0.7550.8300.8500.920

This MS MARCO-tuned L6 model is close to the DistilBERT baseline while being much faster. mentions is best at HotpotQA and 2Wiki recall, while spaCy is best on MuSiQue and produces the highest 2Wiki MRR.

msmarco-MiniLM-L12-cos-v5: 100 questions, shared corpus

Run at commit 0db8dcf with GRAPHMEM_EMBEDDING_MODEL=sentence-transformers/msmarco-MiniLM-L12-cos-v5, the same CUDA release configuration, RTX 4060, batch 16, corpus, graph extraction, and first 100 validation questions as the other runs. The complete sweep (11 reembeds) took 197 s.

datasetgraphmoderecall@2recall@5recall@10MRR
HotpotQAnoneembeddings0.5650.6900.8000.830
HotpotQAnonefts-raw0.5400.7400.8950.872
HotpotQAnonefts-or0.5500.7450.8950.885
HotpotQAmentionsembeddings0.5700.7850.9150.840
HotpotQAspacyembeddings0.5850.7450.8900.848
2Wikinoneembeddings0.5750.6800.7030.967
2Wikinonefts-raw0.6050.6900.7550.950
2Wikinonefts-or0.6050.6900.7550.950
2Wikimentionsembeddings0.7000.9150.9650.983
2Wikispacyembeddings0.6770.8350.9120.995
2Wikioracleembeddings0.8200.9630.9680.995
MuSiQuenoneembeddings0.4750.5450.6000.885
MuSiQuenonefts-raw0.4650.5200.5650.812
MuSiQuenonefts-or0.4650.5200.5650.812
MuSiQuementionsembeddings0.5550.6800.7000.898
MuSiQuespacyembeddings0.5900.7650.8450.899
MuSiQueoracleembeddings0.7650.8450.8650.935

L12 improves selected results over L6 (HotpotQA no-graph recall@2 and MuSiQue MRR), but it is 1.8x slower with no consistent graph-assisted recall gain. L6-cos is the better default for this workload.

Best non-oracle configuration (MS MARCO DistilBERT)

datasetgraphrecall@2recall@5recall@10
HotpotQAmentions0.6450.8250.940
2Wikimentions0.7230.9150.968
MuSiQuespacy0.6050.7350.855

Adding a graph improves every dataset at the top ranks; without one, embeddings are roughly level with BM25 (HotpotQA and 2Wiki FTS even lead at recall@10).

50 vs 100 questions

The ordering above already held at 50 questions; absolute numbers were higher because the harder tail is absent. The 50-question run is also the CPU-parity check (identical recall/MRR to the earlier CPU batch-1 run).

datasetgraph50q recall@2 / @5 / @10100q recall@2 / @5 / @10
HotpotQAmentions0.650 / 0.840 / 0.9300.645 / 0.825 / 0.940
HotpotQAspacy0.670 / 0.840 / 0.9300.620 / 0.800 / 0.920
2Wikimentions0.700 / 0.910 / 0.9650.723 / 0.915 / 0.968
2Wikispacy0.660 / 0.825 / 0.9350.705 / 0.833 / 0.907
MuSiQuementions0.620 / 0.690 / 0.7300.560 / 0.685 / 0.740
MuSiQuespacy0.700 / 0.820 / 0.9100.605 / 0.735 / 0.855

mentions and spacy trade blows: at 100, spacy wins MuSiQue, mentions wins HotpotQA and 2Wiki; the two are within noise of each other.

Lexical recall fix

Before, FTS5 required every whitespace term (implicit AND), so a full natural -language question matched nothing. Now a plain query ORs its terms and BM25 ranks; quoted phrases and prefix* are preserved, and only an uppercase AND / OR / NOT / NEAR switches to strict FTS5 syntax.

datasetfts-raw beforefts-raw nowfts-or
HotpotQA0 / 0 / 0 / 00.540 / 0.740 / 0.895 / 0.8720.550 / 0.745 / 0.895 / 0.885
2Wiki0 / 0 / 0 / 00.605 / 0.690 / 0.755 / 0.9500.605 / 0.690 / 0.755 / 0.950
MuSiQue0 / 0 / 0 / 00.465 / 0.520 / 0.565 / 0.8120.465 / 0.520 / 0.565 / 0.812

Because the OR-join now applies to plain queries, fts-raw is effectively a second OR baseline rather than an all-terms baseline.

Performance

gmem reembed over 100 HotpotQA paragraphs:

backendbatchtime
CPU1~40 s
CPU850.4 s
CPU3278.8 s
WGPU (Vega iGPU)111.2 s (3.7×)
WGPU4 / 8 / 16~11.4–11.7 s
  • CPU defaults to batch 1: padding to the longest text costs more than batching saves there. CUDA defaults to batch 16.
  • The full 100-question sweep (3 datasets, all graphs/modes, 11 reembeds) took 175 s on CUDA vs 72 s for the 50-question sweep.
  • The same 100-question CUDA sweep took 104 s with all-MiniLM-L6-v2.
  • The same sweep took 217 s with all-MiniLM-L12-v2.
  • The same sweep took 112 s with msmarco-MiniLM-L6-cos-v5.
  • The same sweep took 197 s with msmarco-MiniLM-L12-cos-v5.
  • All backends produce identical recall/MRR; the batching changes speed only.

Notes

  • n = 100 questions per dataset; one recall point ≈ one question, so ±0.02–0.03 is noise.
  • spacy extracts fewer than 1 relation per paragraph; paragraphs are linked mostly through shared entities.
  • oracle confirms the ceiling: with gold triples, 2Wiki reaches recall@10 ≥ 0.965 and MuSiQue 0.885, well above the no-LLM graphs.

Pending

  • More questions (or the full validation split) to tighten confidence.
  • Reuse memory vectors across graph runs instead of re-embedding per graph.
  • Tune [retrieval] (damping, memory_seed_weight, …) side by side.
  • Improve relation extraction (broader spaCy patterns or LLM OpenIE), or a combined mentions + spacy graph.
  • Measure RUSTFLAGS="-C target-cpu=native" for CPU builds.