Benchmark engineering for agentic AI

Benchmarks that hold up.

The tasks, environments and audits that tell frontier labs what their agents can actually do.

starzo harness / run 2187 00:00.0 elapsed

Agent trajectory

$ 
Four tasks from four categories, one run each. Every task we deliver proves itself this way before it leaves.

An agent takes the shortest path to reward. If the grader can be reached, it will be reached. If the answer leaks, it will be found. So we build the task, hide the verifier, and attack it ourselves before anyone else can.

What we ship

Three deliverables. One standard.

01

Benchmark tasks

Execution-based, containerised, delivered in your schema with a hidden verifier and a reference solution on every task.

task / claims-adjudication-07Operations
Workspace
starter/ · 14 files
Reference
passes
Empty submission
0.00
Verifier
Hidden
Provenance
author · reviewer · 32 trials

02

RL environments

Reward set by a verifier the policy never sees, hardened against reward hacking, calibrated by repeated trials to the band you order.

env / supply-rerouteBand 0.30 – 0.40
Measured pass rate0.35
Reward-hack sweeps0 paths

03

Evaluation hardening

Your existing suite, attacked the way a capable agent would attack it. Every finding comes with a reproduction and a fix.

audit / suite-v32 findings
  • Answer leakageexpected output committed in tests/fixturesfixed
  • Grader reachableverifier importable from the workspacefixed
  • Reward pathno unintended path foundclear

How every task is built

Taken apart before it ships.

task / newBrief
Scoping0 / 4 fields
Capability
Harness
Passing means
Band
Nothing is written until these four are agreed.
Assembling0 / 3 gates

              
  1. Reference1.00
  2. Empty run0.00
  3. Deterministic32 / 32
Same input, same score0 / 32
Adversarial runs0 / 200
    0 found · fixed0 blocked0 reward paths open
    Measured difficulty—asserted 0.58
    0.25.50.751
    asserted
    0.58measured
    target band 0.35 – 0.45
    Delivered—
    1. 01

      Brief

      We scope the capability you want measured, the harness it must land in, and what counts as passing, before anything is written.

    2. 02

      Build

      Working engineers write the task, the verifier and the reference. The reference passes. An empty submission scores zero. Same input, same score, every run.

    3. 03

      Review

      Our auditors attack it the way a capable agent would: grader access, answer leakage, unintended reward paths. What we find in our own work is what you never see in yours.

    4. 04

      Calibrate and deliver

      Difficulty is measured across repeated trials, never asserted from a read-through. The suite ships into your infrastructure with provenance per task.

    Domains

    Twenty-eight domains, from consensus to counterpoint.

    Long-horizon, execution-based tasks in the fields frontier labs need measured, each written by people who work in the field.

    • 01

      Software engineering

      Application development, debugging, build systems, software tooling

      Software and systems
    • 02

      Distributed systems and networking

      Consensus, concurrency, memory management, congestion control

      Software and systems
    • 03

      Databases and data engineering

      Query optimisation, streaming pipelines, lakehouses, temporal data

      Software and systems
    • 04

      Performance and kernel engineering

      GPU compute, CPU vectorisation, RISC-V, numerical kernels

      Software and systems
    • 05

      Cloud operations, DevOps and SRE

      Kubernetes, incident diagnosis, service reliability, infrastructure troubleshooting

      Software and systems
    • 06

      Software reverse engineering

      Behavioural reconstruction, compatibility, protocol and CLI reimplementation

      Software and systems
    • 07

      Application security

      Secure APIs, denial-of-service resilience, vulnerability repair

      Software and systems
    • 08

      Machine-learning training

      Distributed training, gradient accumulation, reward-model training

      AI and machine learning
    • 09

      LLM inference and serving

      Tokenisation, quantisation, decoding, scheduling, cache management

      AI and machine learning
    • 10

      AI evaluation engineering

      Evaluation harnesses, programmatic grading, judge aggregation, reproducibility

      AI and machine learning
    • 11

      Applied machine learning

      Computer vision, NLP, speech and audio recognition, recommendation systems

      AI and machine learning
    • 12

      Mathematics and formal verification

      Numerical analysis, optimisation, proof certificates, theorem proving

      Science and data
    • 13

      Physics

      Quantum control, trapped-ion systems, gravitational-wave analysis

      Science and data
    • 14

      Chemistry and materials science

      Electrochemistry, battery modelling, spectroscopy, crystallography, kinetics

      Science and data
    • 15

      Biology and bioinformatics

      Genomics, structural variation, genome reconstruction, systems biology

      Science and data
    • 16

      Data science and statistics

      Statistical modelling, hypothesis testing, visualisation, analytical debugging

      Science and data
    • 17

      Economics and econometrics

      Causal inference, panel analysis, policy evaluation, economic data

      Science and data
    • 18

      Digital hardware engineering

      RTL, Verilog, FPGA workflows, arithmetic circuits, communication protocols

      Hardware and CAD
    • 19

      Mechanical engineering and CAD

      Parametric modelling, sheet metal, injection moulding, additive manufacturing

      Hardware and CAD
    • 20

      Finance and fund accounting

      Structured credit, collateral management, valuation, NAV restatement

      Operations and finance
    • 21

      Insurance and claims

      Reinsurance, catastrophe claims, marine insurance, claims adjudication

      Operations and finance
    • 22

      Regulatory and trade compliance

      Sanctions screening, ownership attribution, origin determination, emissions reporting

      Operations and finance
    • 23

      Pharmacovigilance

      Safety-case triage, adverse-event reporting workflows

      Operations and finance
    • 24

      Transportation and logistics

      Rail-yard planning, hazardous cargo, maritime stowage, disruption handling

      Operations and finance
    • 25

      Supply-chain operations

      Cold-chain management, traceability, allocation, sourcing

      Operations and finance
    • 26

      Marketing analytics and advertising technology

      Marketing-mix modelling, incrementality, bidding, budget pacing, reconciliation

      Operations and finance
    • 27

      Visual design and motion graphics

      Typography, font engineering, animation, keyframe reconstruction

      Media
    • 28

      Audio engineering and music

      Signal processing, mastering, music theory, counterpoint

      Media

    Every task is written by someone qualified to grade it.

    Working engineers and domain experts, matched to the category that fits their depth. Not a crowd.