CORE DOMAINS:
AI MEMORY (3) RETRIEVAL & EVALS (4) INFERENCE ARCHITECTURES (6) AGENTS & TOOL CALLING (3) ML SYSTEMS (2) PRODUCT & TASTE (3)
[ LONGMEMEVAL EXPERIMENT ]

When Graph Retrieval Fails to Transfer

A typed relation graph raised synthetic multi-hop recall@10 from 0.325 to 0.800. On a small LongMemEval sample, it fell below BM25. An early research note on benchmark-shaped wins, real dialogue, and what failed next.

0.325 → 0.800
SYNTHETIC GRAPH GAIN
0.168 vs 0.193 (BM25)
REAL DIALOGUE RECALL@10
0.438 (+46% RELATIVE)
COARSE-TO-FINE DENSE
RESEARCH STANDARDS PROOF, NOT CLAIMS
01 / Negative Results are First-Class: If an architecture fails to transfer or collapses on real data, that finding is published and analyzed rather than swept away.
02 / Grounded Provenance: Every benchmark number must trace to specific repository commits, environment configs, and reproducible scripts.
03 / Narrow Scope Precision: Avoid universal sweeping claims about broad technology families (like "GraphRAG is broken"); isolate specific schema, extractor, and retrieval trade-offs.