Bench

Measured. Including the bad numbers.

What this means for you: when a step-written test is green, a real check passed.

  • 0 false passes on step-written tests
  • 0 AI calls on replay
  • 11/11 authored correctly

Bench is an open benchmark for testing tools: small fixture apps, each with a correct build, builds with real bugs and a build where only the look changed. We run our own engine against it and publish what we get, false passes first.

All numbers below are copied from result files committed in the engine repository, with the commit and date they were measured. They are small fixtures, single runs unless a row says otherwise, and not a promise about your app.

Replay: no AI at all

Tests written as exact steps, recorded once, replayed. Measured 30 September 2026 on engine commit f8a19d7, with 10 reruns (bench/baseline.json).

FixtureFalse passWrong failFlakeReplay hitsAI
Web shop (11 tests, 8 builds)0 of 150 of 730 of 11 (10 reruns)58 of 58 steps as recorded0 AI calls
Android app (7 tests, 6 builds)0 of 100 of 320 of 7 (10 reruns)34 of 34 steps as recorded0 AI calls

After a purely cosmetic redesign of the web shop, 8 of 11 tests passed or healed with no AI and no re-recording, and 53 of 56 steps were done without AI. The rest need the fixer model. On the web shop, the replay's verdict and the generated plain-Playwright spec's verdict agreed on all 27 compared tests.

This is where "zero false passes" comes from: tests written as steps, checked by code. We make that claim for step-written tests only.

Writing: how well the AI turns your words into a test

The hosted AI (DeepSeek-V4.1-Flash through OpenRouter, pinned to DeepInfra fp8, zero data retention, no fallbacks), measured 8 October 2026 on engine commit 5ab6cf2 (bench/results/2026-10-08-eval2-corpus*.json).

Each row is one run of the 11 shop tests, written in that style, authored by the model from scratch, then replayed with no AI on the 8 builds. "False pass" is a buggy build where the test passed (15 cases). "Wrong fail" is a correct build where it failed (73 cases).

How the test was writtenInputAuthored correctlyFalse passWrong fail
Tidy numbered stepsa test file11/110/151/73
Mixed prose and stepsa test file11/110/151/73
Given/When/Thena test file10/110/158/73
Terse developer notesa test file9/110/1514/73
Verbose paragrapha test file9/110/1514/73
Spoken transcripta test file9/110/1514/73
Acceptance criteriaa test file8/110/1522/73
Sloppy (typos, wrong labels)a test file7/110/1527/73
Verbose paragrapha description8/110/1521/73
Terse developer notesa description6/110/1534/73

What this says, plainly. The wording you use matters a lot. Tidy steps are authored correctly 11 times out of 11; a terse one-line description only 6 times out of 11, with 34 wrong fails in 73. Wrong fails are the failure mode that costs you time: the test is wrong, not your app, and you fix it in the editor.

What we don't claim. In every row above no test passed on a buggy build, but each row is one run of 15 such cases, which is too small to promise anything for prose or speech. So we claim zero false passes only for step-written tests, and we show the rest as measured.

What it measures

False pass
A test that passed on a build with a real bug. The one number that matters most: a green test that should be red is worse than no test.
Wrong fail
A test that failed on a correct build. Annoying, and it's where AI-written tests lose points.
Flake
A result that changes over 10 reruns of the same build.
Replay hit rate
Steps done exactly as recorded, with no AI, on an unchanged app.
Cosmetic
On a build where the look changed but the behaviour didn't: how many tests still pass or heal without re-recording.

Method

  • Every fixture has a gold manifest: the verdict a correct engine must reach for every test on every build. Results are scored against it.
  • The fixtures, the manifest and the runner are open source (MIT). Run it yourself: optestra bench, or read the CLI reference.
  • We haven't yet run the same suite against other tools. When we do, the method and every trace will be published with the results, whoever wins.