Code-retrieval quality and cost
Profile public-core: 1000 held-out queries across 4 public CoIR tasks, repeated 3 times. No private corpus or local path is included.
Aggregate results
| Mode | nDCG@10 | MRR@10 | P@5 | R@20 | Warm p95 | Index size |
|---|---|---|---|---|---|---|
blended | 0.2625 | 0.2167 | 0.0605 | 0.4683 | 238.82 ms | 91.14 MiB |
neural | 0.2633 | 0.2178 | 0.0608 | 0.4690 | 242.35 ms | 91.14 MiB |
Population variance, phase timings, peak RSS, binary identity, dataset revisions, and checksums are retained in the raw JSON.
Change from frozen baseline
neural improves nDCG@10 by +13.31% and MRR@10 by +13.21% over hash at commit 49b1571de77a. The raw JSON retains every task and run.
Per-task quality
| Task | Mode | nDCG@10 | MRR@10 | R@20 |
|---|---|---|---|---|
codetrans-dl | blended | 0.2259 | 0.1396 | 0.5648 |
codetrans-dl | neural | 0.2290 | 0.1435 | 0.5722 |
codetrans-contest | blended | 0.3777 | 0.3392 | 0.5068 |
codetrans-contest | neural | 0.3778 | 0.3393 | 0.5068 |
cosqa | blended | 0.1501 | 0.1127 | 0.3633 |
cosqa | neural | 0.1505 | 0.1135 | 0.3620 |
codefeedback-st | blended | 0.6394 | 0.6083 | 0.7374 |
codefeedback-st | neural | 0.6394 | 0.6083 | 0.7374 |
Every retained task remains visible so aggregate improvements cannot hide regressions.
Scope
Matrix covers held-out natural-language and code-to-code retrieval. Exact-search tools require a separate exact-query workload.
Artifacts use pinned public datasets and retain source revisions and checksums in raw JSON. Deterministic samples: codefeedback-st (99 queries, 5000 documents). Dataset cards do not declare licenses for: codetrans-dl, codetrans-contest, cosqa, codefeedback-st. Treat downloaded corpora as evaluation inputs; do not redistribute them without checking upstream terms.