Blog

How AI test generation works (and why you can trust it)

Daniel Harkin
Date Updated :
September 3, 2026

Trust in AI-generated code is rising. According to Checksum's own research, 78.1% of engineering leaders trust AI-generated code more than they did a year ago. But in the same survey, 61% admitted they had shipped an AI-originated production incident in the last 90 days, and 74.3% rolled back an AI change that broke in a way their unit tests never caught.

Rising trust without a way to check the work suggests the system is broken.

But is AI testing the solution? Yes, but only if the testing can prove its own findings. It's a concern we hear a lot from prospects, usually posed as a question: did the AI actually find this, or did it memorize the demo?

It's a fair question that deserves a real answer. Here's a breakdown of how our E2E agent works—from how it understands your system and detects the tests to write, through to generating and delivering them—with enough detail that you can judge it for yourself.

Key takeaways

  • Checksum grounds every test in evidence, never invention. Tests are created based on three sources: the application's source code, live exploration of the running app, and a customer's own recorded walkthrough when one exists. The agent isn't allowed to guess a selector or assume a behavior it hasn't observed.
  • Every test passes through implementation, independent review, and verification before it reaches a repo. A dedicated reviewer checks assertion strength and test fidelity, and a separate verification stage confirms the test actually runs and passes for the right reasons. Auto-healing later runs that same discipline in reverse, diagnosing what broke before changing anything.
  • Proof of success is a repeatable pattern. Checksum customer Counterpart catches one significant production issue every week and hasn't had a production outage since adopting the platform; ClearPoint Strategy catches six critical bugs weekly. 

Three ways to detect tests

Checksum draws on up to three sources of evidence, used alone or combined: the application's own source code, live exploration of the running app, and, when one exists, a recording of a real person completing the workflow.

Source-code analysis finds routes, forms, CRUD operations, business workflows, role-based access, and API endpoints, then checks that against what's already tested to find the gaps, not just a list of components to cover. When source access isn't available, live exploration does the equivalent by driving the running app directly and going past the obvious pages: if a list has clickable records, the agent opens at least one to find the detail page, the tabs, and the secondary actions underneath.

The third source is where Checksum purposely draws on a customer's own recorded walkthrough. If it exists, Checksum will treat that recording as the primary source of truth for that specific journey. It won't invent steps the person didn't take, reorder what they did, or quietly expand the flow. Depending on scope, the agent either reproduces the recording exactly as one test, splits it into the distinct goals it contains, or uses it as a seed to find closely related flows nearby. It's never used as an open invitation to improvise.

It’s designed that way because a test that faithfully reproduces what a real user did is doing its job. Independent discovery—finding a flow, a role, or an endpoint nobody walked through in a demo—comes from the other two sources: reading the application's own code and exploring it live. That's where the thing nobody demoed gets found.

Multiple stages, not one prompt

Ask a model to write an end-to-end test and you get one pass: code that might work, might not, and either way reflects a single attempt. That's not what happens with Checksum, though how many stages a given test passes through depends on what's being asked of it.

For a bounded, well-defined request, generation runs a lean path. The planning still happens, it's just handled inline by the implementation agent's own reasoning rather than broken out as a separate step.

For broader work, anything spanning many stories, ambiguous requirements, or changes to shared utilities, planning becomes its own explicit stage first. It looks at which files and utilities are involved, which role should run the flow, where setup can lean on an API instead of the UI, and what each assertion needs to prove before a line of test code gets written. More is at stake, so more reasoning happens before anything gets built.

From there, every request runs through the same stages:

  1. Implement: Write the actual Playwright code, grounded in evidence, since the agent isn't allowed to guess a selector or invent behavior that wasn't observed.
  2. Review: A separate pass, run by a dedicated reviewer, that checks assertion strength alongside selector quality, test isolation, and whether the generated test still matches the story it came from. An assertion that accepts two possible outcomes instead of the story specified is exactly what this stage exists to catch, before the test ever reaches your repo.
  3. Normalize: the generated code is converted into Checksum's own execution format, correct assertion placement and imports included, so it runs the same way every time.
  4. Verify: Confirm the test starts, authentication works, every selector resolves, every assertion passes, and the resulting flow matches the story it was built from. If verification turns up a real bug in the application, Checksum won’t weaken the test until it goes green. The bug stays visible instead.

"Does it click through the app, or just hit the URL?"

We get asked this a lot and the answer is both, but not for the same thing. When a generated test needs to reach a starting state, logged in as the right role, with the right records already in place, Checksum prefers an API call if the codebase already has one to reuse. For the sake of speed and reliability, nobody needs a test to spend three minutes proving a login form works before it gets to the thing actually being tested.

The journey itself is addressed differently. Every action in a generated test has to be grounded in evidence: a recorded click, a route found through live exploration, a form discovered in the application code. Locators are chosen in a deliberate order: test IDs first, then accessible roles and labels, then stable visible text, and a CSS selector only when nothing else is available. Positional selectors and deep DOM paths used to dodge ambiguity are explicitly discouraged.

None of that is what you'd build if the goal were to shortcut past the interface. It's designed this way so that tests keep working after the interface inevitably changes.

What proof looks like

Organizations that use Checksum don't measure the platform’s impact using one-off headline metrics. Rather, success comes from the same detection sources doing what they're built to do, run after run.

Counterpart uses Checksum to catch one significant production issue per week and the company hasn’t experienced a production outage since adopting the platform. ClearPoint Strategy catches six critical bugs weekly; that’s a standing pattern rather than a one-time demo result.

Autonomous doesn't mean unaccountable

Every test Checksum produces still lands as a pull request in a repository you own: standard Playwright code, reviewable line by line with no lock-in. But the accountability doesn't stop at delivery.

When a test breaks later because the app changed, Checksum's auto-healing runs the same discipline in reverse, and it starts by working out what actually went wrong before it touches anything. Is this an outdated test, a real product bug, bad infrastructure, or something an engineer should decide? Whatever isn't a genuine test problem gets left alone. Checksum won't quietly patch a test into looking fine when the app underneath it is actually broken.

What's left gets repaired, but under real constraints: keep what the test was originally checking, change as little as possible, and fix the step that actually failed. Patching around it to force a green run doesn't count as a fix. And the shortcuts that paper over instability—forcing a click through, grabbing the first matching element instead of the right one, throwing in a wait-and-hope pause—are treated as band-aids, not fixes.

Then another agent checks the work. A second, independent reviewer, with no stake in the fix it is grading, runs it against nine separate criteria: did it preserve the original intent, is it the minimal change, did it correctly flag a real bug instead of quietly working around it? If it doesn't hold up, it goes back for another pass. Only once it survives that does the fix run, for real, before becoming a PR.

That's what "autonomous" means inside Checksum's own process, from the first test it ever writes to the tenth time it repairs one. It's never just one model deciding it's finished, but a second, independent one whose only job is to disagree with it if it's wrong.

See it for yourself

The data says trust in AI-generated work is outpacing the industry's ability to verify it—in the code itself and in the tests meant to catch what breaks. CI/CD automated how code gets delivered. Continuous verification is supposed to automate how it gets proven, and that only becomes meaningful when you see the proof.

Point Checksum's E2E Agent at your own application and watch what it finds, or read the full State of AI code report for the data behind why this matters right now.

FAQs

How does Checksum know its AI-generated tests aren't just copying a demo?
Checksum draws on up to three evidence sources (the application's source code, live exploration of the running app, and a customer's recorded walkthrough when possible) and grounds every action in what it actually finds. Independent discovery of flows nobody demoed comes specifically from reading the code and exploring the live app, not from the recording.

Does Checksum's AI just click through the UI, or does it call APIs directly?
It does both for different purposes. To reach a starting state (i.e. logged in as the right role, with the right data in place) Checksum prefers a fast, reliable API call if one already exists in the codebase. The actual journey being tested is still grounded in evidence: a recorded click, a route found through live exploration, or a form discovered in the application code.

What happens if Checksum's auto-healing can't tell whether a test failure is a bug or an outdated test?
Checksum classifies every failure before touching anything: an outdated test, a real product bug, bad infrastructure, or something that needs a person to look at it. If the failure can't be resolved automatically, it's routed to that last category and left for a human to review, rather than guessed at. Checksum will never patch a test to force it green when the uncertainty is really about the application underneath. Every repair that does go ahead also passes through a second, independent review before becoming a PR.

Daniel Harkin

Daniel Harkin is Senior Software Architect at Checksum, an AI-first company solving one of the most critical challenges in modern software development: ensuring the quality of code that ships. Daniel brings over 23 years of experience building and scaling complex software systems across fintech and adtech, enterprise and startup. Most recently, he spent nearly a decade as Senior Software Engineering Manager at Quantcast, where he led engineering teams building out high-scale data processing and advertising technology. Prior to that, he held senior engineering roles at Goldman Sachs, IG Group, and J.P. Morgan.