Most test tools give you two answers. Green or red. Pass or fail.
That was fine when a person ran the tests and read the output. A person can see “the browser did not start” and know that this is not the same as “the feature is broken.”
An AI agent running a loop at 2 a.m. cannot see that difference unless you give it one.
I run loops like that because I build a product with a small team and a big scope. Agents run improvement loops on the model and the app while we sleep. That only works if I can trust what the loop tells me in the morning. So I give every check in my projects three outcomes:
- pass: the gate is clear. Go to the next step.
- fail: a tracked metric got worse. Revert the change and try again.
- error: the check could not run. Stop and alert a person. It never counts as a fail or a pass.
This post explains why the third answer matters, the audit that taught me this, and the small contract I now use everywhere.
Self-improvement loops, Karpathy style
In March 2026, Andrej Karpathy released autoresearch. An agent edits a training script, trains for five minutes, and checks one number. If the number got better, it keeps the change. If not, it resets with git and tries the next idea. That is about 12 experiments an hour, or about 100 while you sleep. The instructions tell the agent never to stop and ask, because the human “might be asleep.”
I run my detector improvement loops in the same shape: one change, one run with a fixed budget, one score against a baseline, then keep or revert.
Look at the autoresearch log and you see three states already. Every row in results.tsv is keep, discard or crash. Even the simplest self-improvement loop needs a third word, because a run that crashed measured nothing.
But “my change crashed” and “the harness could not run” are not the same thing.
In autoresearch, crash means the idea broke: a typo, a model too big for memory, an idea that does not work. The right move is to fix it or skip it, and continue. That is close to fail. The change was bad.
An error is different. The GPU is gone. The data path is wrong. The eval wrote nothing. Every next idea will hit the same wall. A loop that is told never to stop will spend the whole night there, and log 100 “crashes” that are really one problem. So my loops keep the two apart. A crash that the change caused counts against the change. An error in the harness stops the loop and calls a person.
Two ways to get it wrong
If you only have two answers, an error must become one of them. Both choices hurt.
Error counted as fail: the loop thrashes. Say the GPU runs out of memory during an eval. The loop sees “fail,” reverts a good change, and tries something else. The next run hits the same memory limit. Now the agent “learns” that every change is bad. It burns hours and throws away good work, because the problem was never the change.
Error counted as pass: the check lies. This one is worse, because it is quiet. A browser test crashes before it writes its result. A test file cannot import a dependency. The runner exits, nothing is red, and the pipeline moves on. Your dashboard says green. Nothing was tested.
I call this “a check that lies about having run.” With people in the loop, someone usually notices within a day. With agents in the loop, nobody looks. The agent trusts the green and builds on it.
The audit that taught me this
I build a construction-tech product that reads HVAC drawings. Its main release gate, one slice of the product’s ~1,100 automated tests, had five layers: build and type check, 358 unit tests, server parity tests, a browser smoke test with 16 checks against stubbed backends, and an API suite of 77 tests in Docker.
On paper, that is a strong gate. Then I audited every testing surface in the repo and actually ran them. I found two problems on the first day.
1. The gate was red on the main branch. A parity test compares two copies of an OCR server: the one on the GPU machine and the one behind the API. The GPU copy had changed five commits earlier. The API copy had not. The parity test did its job. It failed correctly. But the pre-push hook that runs it was opt-in, and it was not installed on the machine that pushed. So the drift landed anyway.
I was annoyed, mostly at myself. I wrote that hook, and I made it opt-in. My own shortcut let the drift through.
2. A whole test layer could not run, and nobody knew. One test file could not import a package. The package was declared in the project, but the local install was stale. The test layer did not fail. It just did not run. After one install command, all 358 tests passed.
Both problems had the same root cause:
Verification existed, but nothing forced it to run.
Neither was a “fail.” Both were “error”: the check could not do its job. And in both cases, the system treated that as silence, and silence looked like a pass.
What it really cost me was trust. After the audit, I stopped believing green. Every pass needed a second look, and that slowed the whole team down until the gate could tell the truth again.
The contract: one small JSON verdict
Now every gate, eval and probe writes the same shape:
{
"status": "fail",
"blocking": ["recall.return_grille dropped 0.07 vs baseline (limit 0.05)"],
"metrics": {
"precision": 0.91,
"recall": 0.78,
"count_error": "+6%"
},
"confidence": 1.0
}
statusis exactly one ofpass,failorerror.blockinglists the reasons in plain words. It is empty on a pass.metricsholds raw numbers, so a person or an agent can see how bad it is.confidenceis 1.0 for deterministic checks. It is lower only when a model is the judge.
It is a JSON Schema file in the repo. Anything that reads verdicts validates against it.
The verdict logic is pure, and it has its own tests
The rules that decide “may this run call itself green?” are the most dangerous code in the harness. If they are wrong, every other check is decoration. So I moved them out of the scripts that start browsers and call servers, into small pure functions with no I/O. Then I wrote tests for them.
Here is an illustrative version:
// Pure: callers pass in what they already read. No files, no network.
export function classify({ verdictText, exitCode, ranThisRun }) {
// No verdict from THIS run is an error, even if an old one exists.
if (!ranThisRun || verdictText == null) return "error";
let status;
try {
status = JSON.parse(verdictText).status;
} catch {
return "error"; // unreadable verdict: the check did not finish
}
if (!["pass", "fail", "error"].includes(status)) return "error";
// A child that said "pass" but exited non-zero contradicted itself.
if (status === "pass" && exitCode !== 0) return "error";
return status;
}
// A gate that can never fire is not a gate.
export function deadGates(baseline, gatedClasses, maxDrop) {
return gatedClasses.filter((c) => {
const b = baseline?.perClass?.[c];
return b && b.truth > 0 && b.recall - maxDrop < 0;
});
}
Two rules in there came from real bugs:
- Stale verdicts. The browser test wrote its result to a file. The folder was ignored by git, so the file survived between runs. When the browser layer died, the previous run’s “pass” was still there, and the gate read it. Now the gate deletes the file before it starts the test, and a missing file is
error. - Dead gates. One baseline was saved from a run where the model predicted nothing. The rule was “fail if recall drops more than 0.10 below baseline.” With a baseline recall of zero, that rule can never fire. But the log still printed “gated within 0.10 of baseline.” The check reported enforcement it did not perform. Now the harness detects that case and says so.
How loops and agents use the verdict
The verdict is the interface between the harness and whatever runs it.
- An improve loop reads
status. Onpass, it keeps the change and records the numbers. Onfail, it reverts and tries the next idea. Onerror, it stops and writes why. - A Stop hook on an AI coding agent reads the verdict before the agent may end its turn. If the status is not
pass, the agent cannot claim it is done. - A person reads
blockingandmetrics. One line tells them what happened. No log digging.
The same idea works for plain CI. On a team marketplace app I work on, one command runs the backend, web and mobile checks. Agents run the JSON mode of that command:
$ make verify-json
{"status":"fail","blocking":["frontend: 2 type errors"],
"metrics":{"backend":"pass","frontend":"fail","mobile":"pass"}}
Agents do not parse test output. They read one object. The repo rules say: run verify before every push, and never merge red.
What the agent starts with
autoresearch gives the agent one instruction file, program.md. I use the same convention. Before the first run, the agent reads one file that says what it may know, what it may touch and when it must stop. I learned that the loop is only as good as that file.
This is what my detector loop’s program.md gives the agent:
| Input | What it contains | Why it matters |
|---|---|---|
| Scope | One file it may edit: the training script. Data prep and evaluation are fixed. No new packages. | The agent cannot “improve” the score by changing the test. |
| Domain facts | 26 symbol classes, a small dataset, heavy class imbalance, black-and-white drawings that are always axis-aligned, classes that look the same without context. | It skips ideas that cannot help, such as color augmentation or rotation. |
| Hardware | One remote GPU, its memory limit, how to sync, run and read the log. | Out of memory becomes a known crash. |
| Goal and tie-breakers | Maximize the main score. No class may drop to zero. Rare classes count more. Simpler wins a tie. | A higher average that loses a class is a regression. |
| Search plan | Phase 1: change one variable at a time. Phase 2: combine the winners. Phase 3: fine-tune near the best. Plus a ranked list of the knobs that matter for this data. | The agent explores in order, not at random. |
| Memory | Every run goes into a database: full config, all metrics, per-class scores and the exact training source. A suggest command reads it. | “Check the history before you change anything” stops repeated failed ideas. |
| Stop rules | Never stop for pass or fail. Stop on error. | The night goes to experiments, not to a broken setup. |
What the agent does not get matters as much. The gate’s fixtures are synthetic. No customer drawing enters the agent’s context. The gate reports exit codes and counts only.
Refining the context, run by run
The first version of the file was short: edit, train, compare, keep or revert. Each lesson since then became a rule in the file.
- Real runs are long. A full run takes 45 to 90 minutes. An agent with a short timeout calls a slow run a crash. Rule: a run has crashed only if the process died or the log stopped moving.
- An average can hide a loss. The overall score can go up while one class goes to zero. Rule: after every keep, check the per-class scores. If a class that worked before now scores zero, discard.
- Agents forget. Without memory, a loop tries the same idea on different nights. Rule: read the experiment history before every change.
- Context goes stale. We moved to a GPU with twice the memory, and the old limit in the file made the agent too careful. Rule: the file starts with a dated status line, and I update it like code.
- The goal changes. When the focus moved from training the model to improving the product, I wrote a second loop with its own
program.md. It improves the product against a fixed set of scorecards. Its rules:- The gate is fixed. The agent never edits it to make it pass.
- Pick the worst problem by user impact, not the metric that is easiest to move. A broken feature beats a weak model.
- Update baselines only with a person’s approval, as a new dated file.
- Write every attempt, with before-and-after numbers, in a journal.
erroris notfail. If the GPU machine is unreachable or a dependency is missing, stop and report. Do not iterate.- If the conclusion is “this needs more training data,” write that down and stop.
So I treat the agent’s context as part of the harness. I version it, date it and change it when a run teaches me something.
Measure what the user gets
The three-valued verdict is half of the lesson. The other half is what you measure.
My detector training reported tile-level mAP. That number answers “did training converge?” It does not answer “when an estimator clicks AI Assist on a full page, how much does it find, and how much junk does it add?”
So the product eval calls the live detection endpoint on full drawing pages and grades each symbol class: precision, recall, F1 and count error. A model can improve on tiles and get worse on pages. Only the page-level eval can say fail for the thing users actually see.
The same applies to matching. In a near-duplicate identity problem, the eval measures correctness: one frozen corpus, one scoring function, every strategy scored on the same rows. It does not measure how many rows were matched, because a wrong match is worse than no match.
Check every link that can break
My rule now: if a moment, a link in the chain or a step in a process can break, something must check it.
Over time, working with agents and loops, you get a feel for where an agent will “cheat” or “forget.” It skips an eval. It reads an old result. It says a run is done when the run never started. Those are the places where the harness needs a check.
That gives you a deeper understanding of how agents work, and of how to prepare your infrastructure to handle them.
A checklist you can apply this week
- Give every check three outcomes. Write down what
errormeans for each one. - Never let a missing result read as a pass. Delete old result files before each run.
- Treat “said pass, exited non-zero” as
error. - Put verdict rules in pure functions, and unit-test them.
- Detect dead gates: thresholds that can never fire with the current baseline.
- Emit one JSON object per run, with
status,blockingand rawmetrics. - Make hooks mandatory, not opt-in. A check nobody runs is not a check.
- Grade the product, not only the training job.
- Give the agent one instruction file: scope, domain facts, goal, search plan, memory and stop rules. Date it, and add a rule after each lesson.
FAQ
What is a three-valued verdict in testing? It is a check result with three states: pass (the gate is clear), fail (a tracked metric regressed, so revert) and error (the check could not run, so stop and alert a person). The third state stops infrastructure problems from being read as code problems or as success.
Why should an error not count as a test failure? Because the change under test was never measured. If an agent loop treats an error as a fail, it reverts good changes and repeats the same broken run. Errors need a person or a fix to the harness, not a new attempt.
What should an AI agent loop do when a check returns error? Stop the loop, record the blocking reason and alert a person. It should not retry blindly, revert the change or continue as if the check passed.
How is this different from Karpathy’s autoresearch loop? autoresearch logs each run as keep, discard or crash, and a crash usually means the change itself broke. That maps to fail. The extra state here is error: the harness itself could not run, so every next attempt would break the same way. A loop that never stops should halt on error instead of logging the same crash all night.
What should a program.md for a self-improvement loop contain? The one file the agent may edit and the files that are fixed, the domain facts that rule out useless ideas, the hardware limits, the goal with its tie-breakers, a search plan, where the experiment history lives, and the stop rules. Add a dated status line and update the file when a run teaches you something.
How do you stop a test from passing when it did not run? Make “no result from this run” an error. Delete old result files before the run, check the exit code against the reported status, validate the result against a schema, and unit-test that logic.