Josh SolisJS Book an intro callBook a call

LLM eval harnesses · AI agent evaluation · CI gates

Harness Engineering: LLM eval harnesses and AI agent evaluation with pass / fail / error verdicts

AI agents write code fast. A fast change that nobody checks is a fast regression. I needed proof for each change to code, to models and to data pipelines.

How it works

Speed is easy. Proof is hard.

Many AI agents work at the same time. Each agent gets one task and its own git worktree.

No change merges on trust. Each change must pass a harness that grades it.

The same idea checks code, models and data pipelines. In ten months, 5,300+ commits went through it.

Figure: dots leave one engineer, split into eight parallel agent lanes, and pass a gate (build, tests, smoke) before they merge. Some dots stop at the gate.

Interactive

Six LLM eval harnesses, one contract

Each harness grades one kind of AI work: code, models, matching or agents. Select a harness to see what it checks and an example verdict.

Every change, before push

Code gate

Five layers run in order: build and type check, unit tests, server parity tests, a browser smoke test with stubbed backends, and an API test suite in Docker.

Why: Agents push many small changes. Each change must prove that nothing else broke.

{
  "status": "pass",
  "blocking": [],
  "metrics": {
    "build": "ok",
    "unit": "358/358",
    "parity": "ok",
    "smoke": "16/16",
    "api": "77/77"
  }
}

Example verdict. The shape is real. Some values are examples or redacted.

The contract

Three answers

Every harness returns the same JSON verdict. A loop, a hook or an agent reads it and acts.

  • pass

    The gate is clear. Go to the next step.

  • fail

    A tracked metric got worse. Revert and try again.

  • error

    The check could not run. Stop and alert a person. This is not a fail.

Video · 38 seconds

Many agents, one gate

  1. One task per worktree.
  2. Every branch meets the gate.
  3. Merge, ship, learn.
Transcript
  1. One developer starts many AI agents.
  2. Each agent works on one branch in its own worktree.
  3. A task must pass every gate.
  4. A failed task goes back to its agent.
  5. Passed tasks merge into main.
  6. The loop repeats every day.
  7. Review and tests set the speed, not typing.

What I built

The parts

  • A five-layer code gate: build, unit tests, parity tests, browser smoke tests, API tests. Agents must pass it before they push.
  • One verify command. It runs the same checks on a laptop and in CI, and prints one JSON summary for agents. Main does not merge without it.
  • Model evals that grade the live product, per class, on full pages. Not only the training metric.
Show 7 more
  • An identity eval: one frozen corpus, one scoring function. Every matching strategy is scored on the same rows.
  • A/B harnesses that score competing recognizers against frozen baselines.
  • End-to-end journey probes: real client code, a real browser and known-answer fixtures.
  • An autonomous training loop: check a run, read the result, pick the next experiment, launch it.
  • I treat a hand-off as a claim to check. My pickup tool checks each claim against the live repo before an agent resumes.
  • Walk-and-work: voice control of coding agents from a phone, with spoken status back.
  • Agent rules files, session hand-off and pickup skills, and 80+ isolated worktrees, so many agents can use the harnesses at once.

Results

By the numbers

6
Harness types
~1,100
Automated tests (one product)
5,300+
Commits / 10 mo

Lesson

An audit found two silent failures. In both, the checks existed but nothing forced them to run. Now the gate runs itself, and a check that cannot run says so.