Skip to content

Benchmarks & Evaluation

Build reproducible benchmark inputs, prepare ground truth for source-location retrieval evaluation, and carry the configuration and provenance needed to interpret every result.

  • Collect

    SWE-bench instances

    Sample representative tasks across repositories and languages without losing their source metadata.

    Collect benchmark inputs →

  • Synthesize

    Query datasets

    Generate traversal, behavioral, and multi-step queries with explicit curation and verification stages.

    Run the synthesis pipeline →

  • Prepare

    Ground-truth locations

    Extract expected files, symbols, and one-based source ranges from benchmark patches.

    Use the GT locator →

  • Package

    Artifact bundles

    Preserve manifests, metrics, diagnostics, and provenance as one verifiable evaluation output.

    Build an artifact bundle →