🧠MoC-RAG Benchmark Leaderboard
Does routed, typed context (Mixture-of-Contexts RAG) beat flat RAG for
agentic memory? This leaderboard reports retrieval, context-efficiency, routing,
and robustness metrics for BM25 / dense / hybrid / metadata-filtered / reranked
RAG and MoC-RAG (top_experts ∈ {1,2,3,all}).
Headline (sentence-transformers): BM25 collapses on adversarial queries (−36% vs keyword), while MoC-RAG holds and overtakes BM25 by ~+15 points on the adversarial split, carrying roughly half the hard distractors of dense RAG at 95–100% routing accuracy.
Dataset: ruslanmv/moc-rag-benchmark.
Embedder
Robustness by query type (the key result)
metadata_rag | 100% | 81% | 64% | 27 | 104 | 103 | -36% |
Single-split leaderboard
metadata_rag | 78% | 19% | 0.7192 | 0.7271 | 167 | 23560 | 12% | 3.9686 | 100% | 83% |
bm25_rag | 78% | 19% | 0.7192 | 0.7271 | 167 | 23560 | 12% | 3.9686 | - | 83% |
dense_rag | 78% | 19% | 0.5582 | 0.6166 | 271 | 23539 | 12% | 3.9721 | - | 83% |
hybrid_rag | 82% | 21% | 0.7232 | 0.7403 | 244 | 23519 | 12% | 4.2094 | - | 84% |
metadata_rag | 82% | 20% | 0.7031 | 0.7214 | 273 | 23364 | 13% | 4.1945 | - | 84% |
reranked_rag | 82% | 21% | 0.7171 | 0.7446 | 249 | 23499 | 12% | 4.2129 | - | 86% |
moc_rag_e1 | 81% | 20% | 0.7484 | 0.7598 | 104 | 23509 | 12% | 4.1473 | 84% | 77% |
moc_rag_e2 | 85% | 21% | 0.7677 | 0.7811 | 122 | 23547 | 13% | 4.3318 | 95% | 86% |
moc_rag_e3 | 86% | 22% | 0.7889 | 0.7963 | 156 | 23529 | 13% | 4.3988 | 99% | 89% |
moc_rag_all | 85% | 21% | 0.7661 | 0.7788 | 196 | 23568 | 12% | 4.3067 | 100% | 89% |