[ LONGMEMEVAL EXPERIMENT ]
When Graph Retrieval Fails to Transfer
A typed relation graph raised synthetic multi-hop recall@10 from 0.325 to 0.800. On a small LongMemEval sample, it fell below BM25. An early research note on benchmark-shaped wins, real dialogue, and what failed next.
0.325 → 0.800
SYNTHETIC GRAPH GAIN
0.168 vs 0.193 (BM25)
REAL DIALOGUE RECALL@10
0.438 (+46% RELATIVE)
COARSE-TO-FINE DENSE