Benchmarks & Evaluation¶
Build reproducible benchmark inputs, prepare ground truth for source-location retrieval evaluation, and carry the configuration and provenance needed to interpret every result.
-
Collect
SWE-bench instances
Sample representative tasks across repositories and languages without losing their source metadata.
-
Synthesize
Query datasets
Generate traversal, behavioral, and multi-step queries with explicit curation and verification stages.
-
Prepare
Ground-truth locations
Extract expected files, symbols, and one-based source ranges from benchmark patches.
-
Package
Artifact bundles
Preserve manifests, metrics, diagnostics, and provenance as one verifiable evaluation output.