New research

The state of
code verification

Benchmarks for manual vs. AI testing

It matters what we measure.

Web agent benchmarks tell us that AI isn't ready for production automation. Popular evaluations report low E2E success rates and imply that agent-driven workflows are unreliable. Yet real companies are already using AI to automate meaningful, recurring web workflows at scale.

This contradiction exists because standard benchmarks measure model behavior in isolation, while production automation depends on resilience over time.

The benchmarks typically used by engineering teams focus on one-shot task completion and raw action accuracy. But real-world automation is different. It involves:

  • Workflows that span dozens of steps
  • Constant UI changes
  • Failures that must be detected, diagnosed, and repaired
  • Retry logic, validation, and fallbacks
  • Deterministic code that remains the source of truth

Reliability is not about whether an agent succeeds once. It's about whether a workflow stays operational over weeks and months, ultimately for as long as the web page and platform exists.

What actually breaks

To understand where tests fail, Checksum analyzed more than one million production automation runs across hundreds of applications.

  • 32% selector changes
  • 27% flow changes
  • 22% environment instability
  • 19% loading and timing issues

Meanwhile, test and automation maintenance is quietly consuming engineering time. Releases are delayed, trust in automation erodes, and coverage stagnates or declines.

Keep reading to see how a suite of 500 tests can cost $4.3M+ annually to maintain, and the difference that AI-powered continuous verification can make to maintaining your test suite.

$4.3M
Annual cost to manually
maintain a large test suite
PART 1

The problem with web agent benchmarks (and why we need better ones)

Why existing web agent benchmarks report misleadingly low success rates and what changes when we track the thing that predicts production success.

When current web agent benchmarks tell us AI isn't ready to automate the web, they're not telling the whole story. Benchmarks like Browserbase's Stagehand evals and agent-evals have fundamental gaps that make them poor predictors of real-world automation success.

The compound action problem

Let's consider Stagehand's atomic action evals. They report around 85% accuracy per action, which sounds pretty good until you do the math on a real workflow.

Booking a table at a restaurant takes maybe 10 actions: navigate to the site, click reservations, select date, select time, enter party size, fill in contact info, submit. Even simple enterprise workflows are longer. At 85% per-action accuracy, your success rate for a 10-action flow is 0.8510 = 19.7%.

One in five attempts succeeds.

The agent-evals benchmark from Stagehand shows end-to-end success rates around 65% on WebVoyager-style tasks. That's more realistic for multi-step flows, but it's still telling us the same thing: you cannot reliably automate web processes with AI today.

Except that people do. So what's missing?

The reliability question

A 65% success rate means different things depending on distribution.

Having 65% of tasks succeed reliably is transformative. If you could automate 65 workflows in a set of 100, most people would happily handle the remaining 35 manual tasks. But if every task succeeds 65% of the time, you can't trust any individual automation to work. The constant checking, retrying, and fixing makes the system unusable.

Current benchmarks don't tell us which world we're in. By giving aggregate numbers without showing the distribution of successes across tasks, the most important question remains unanswerable: which processes can I actually automate today?

The agent harness gap

These models aren't being run raw in production. Instead, every real-world AI automation system has an agent harness: retry logic, validation checks, caching for known-good paths, fallback strategies, and confidence thresholds.

A model that scores 60% raw might hit 90% with good harness logic. Another model at 65% raw might plateau at 70% because it fails in ways that are harder to recover from. Today's benchmark numbers don't tell you which is which.

This matters because different models fail differently. One might make locator errors that are easy to detect and retry. Another might succeed at actions but misunderstand intent in ways that are harder to catch. Raw accuracy scores hide these differences.

Real production systems layer multiple recovery strategies. When Playwright code fails mid-execution, AI can kick in to complete the action in real time. When a selector breaks, the system can try alternative approaches before giving up. When validation fails, it can retry with different parameters. These engineering actions change the success rates dramatically.

The problem is that when benchmarks measure model capability in isolation, they miss that production is about model capability plus engineering.

This isn't just a benchmark problem. In Checksum's survey of 105 engineering leaders for the State of AI code report, AI writes code on every team surveyed, but verification hasn't kept pace:

48.6%
generate AI-assisted
unit tests
69.5%
test under production-like
conditions

The gap between how confidently teams write AI code and how rigorously they verify it is the same gap the benchmarks miss. Capability without a harness doesn't just underperform on an eval—it ships incidents.

The auto-healing problem

Writing the initial browser automation script is the easy part.

A capable engineer can write a Selenium script to automate a web workflow in an afternoon, no AI needed. Spending 5–10 hours on the script for a high-value process that runs 100 times a day is worth it. Spending 30 minutes reviewing AI-generated Playwright code is even easier.

The killer is maintenance. That script breaks next week when the button is redesigned. It breaks again when a loading spinner is added. Again when the form validation changes. And so on.

This is why LLMs matter for browser automation. Not because they can take a vague prompt and figure out a website they've never seen, but because they can adapt when things break.

Today's benchmarks miss that because they don't test healing at the right layer. The most effective approach isn't pure AI that figures everything out from scratch every time. It's AI that maintains deterministic code.

What we really need

The enterprise automation problem is different from the benchmarks. It's not "can AI figure out any website from scratch?" but instead "can AI help me maintain automations that I care about?"

Real enterprise automation looks like this:

  • The engineer defines a repeatable workflow (maybe by reviewing AI-generated Playwright code)
  • The engineer adds structure, guardrails, and validation as needed
  • The automation runs daily, weekly, constantly
  • When code breaks mid-run, AI recovery kicks in to complete the task
  • After runs complete, the system reviews failures and submits fixes
  • Sites change, the automation adapts
  • When it breaks, the engineer gets clear, actionable reports e.g. "this selector broke, here's the fix"
  • Over time, the automation learns the stable paths and gets faster

Taking this approach assumes some human oversight from the beginning. It assumes retry logic and validation. It assumes you care about the same workflows over time, not one-shot tasks. It assumes code as infrastructure, with AI as the maintenance layer—not AI as a black box.

Benchmarks that test one-shot performance on diverse tasks can't measure the thing that matters most for production automation: resilience over time with human-reviewable artifacts. They don't test: "This workflow ran yesterday. Today the UI changed. Can you still complete it?" That's the actual problem that AI needs to solve.

The path forward

Current benchmarks tell us AI isn't ready. But AI is already automating real work for real companies and the gap reveals that we're measuring the wrong things.

Benchmarks need to:

  1. Report per-task success rates: Go beyond aggregates to show which tasks are automatable today.
  2. Test with agent harnesses: Measure production systems with real-time recovery, not lab conditions.
  3. Measure adaptation over time: Run workflows, change the sites, and see what survives. This measures whether AI can maintain Playwright code, not just execute actions from scratch.
  4. Value reliability over breadth: Build on the principle that a system that automates 10 workflows at 95% reliability is more useful than one that attempts 100 at 50%.
  5. Test code generation and healing: Answer if the system can produce reviewable artifacts and fix its own code when things break.

At Checksum, we constantly track benchmarks in the form of our own customer metrics. Not because Stagehand's evals are wrong—they're measuring what they're meant to measure. But because the gap we're measuring is different.

The web is already being automated by AI. The question isn't "can it work?" It's "how do we know which workflows will work, and how do we make them work better?" Better benchmark and metric tracking is the first step.

PART 2

What breaks and how often

When current web agent benchmarks tell us AI isn't ready to automate the web, they're not telling the whole story. Benchmarks like Browserbase's Stagehand evals and agent-evals have fundamental gaps that make them poor predictors of real-world automation success.

We used real customer data and analyzed over one million end-to-end test runs across hundreds of production web applications. These tests run in CI, pre-deploy pipelines, and synthetic production monitoring. The results show what actually causes automation to fail, how often it happens, and what it takes to fix it.

Scope: Tests, but really all web automation

The data in this report comes from UI test automation based on Playwright and Cypress. In practice, the failure patterns apply to almost any web automation: RPA scripts, scraping workflows, monitoring tools, internal bots, and even AI agents executing browser actions.

The same selectors break. The same timing issues show up. The same DOM changes cause failures.

If anything, tests are strictly harder than general automation. They must be deterministic enough for CI. They cannot have a human in the loop when they break. They run in the harshest environments: fresh deployments, feature flags, and partial rollouts.

If AI can keep tests healthy in this environment, it can maintain almost any web automation.

Failure rates in production

How often do automated workflows break in production? Across all customers we measured failures at the test-run level, separating suites that use Checksum from suites that do not.

A typical Checksum team running 500 tests every commit will see an occasional failing test. Teams without Checksum see several failures on almost every CI run. Note: the remainder of this report uses the median percentile for all calculations.

MANUAL/LEGACY
AI-MAINTAINED
(CHECKSUM)

Failure rates by test complexity

Tests are categorized by the number of distinct user actions and pages involved. We found that longer tests accumulate more failure points; a test that touches login, navigation, data entry, and checkout has roughly three times the failure rate of a test that validates a single form submission.

Test complexity Actions Median failures per 100 runs
Simple1–59.3
Moderate6–1514.1
End-to-end journeys16–3021.7
End-to-end journeys30+31.4

Why automation breaks

Looking at 18,000 randomly sampled failures, we tagged each one by primary cause. While multiple things may be wrong in a single run, categories are assigned based on what needed to change for the test to pass again.

Investigating root cause distribution

1. Selector changes: 32% of failures

This category encompasses the classic automation problems: things like a button moving or label changing, renaming a CSS class, or removing an ID. Selectors remain the single largest source of test breakage. The intent of the step is usually unchanged, but the locator is now wrong.

Examples from real production systems:

  • A SaaS onboarding flow where the "Continue" button became "Next"
  • A checkout page where .primary-cta became .Button_buttonPrimary__3fK2x after a design system upgrade
  • A table where #users-table became data-testid="users-table"

Breakdown by selector type

Selector type Share of selector failures Median fix time (human)
CSS class changes41%35 min
ID changes23%25 min
Text / label changes19%20 min
Structural (nth-child, hierarchy)12%55 min
Attribute changes5%30 min

CSS class selectors fail most often because teams frequently use generated class names from CSS-in-JS libraries or design system migrations. Structural selectors (those relying on DOM hierarchy) take longest to fix because they often require understanding why the structure changed.

2. Flow changes: 27% of failures

Flow changes mean the happy path through the product is different from what the test expects. These failures are about intent, not just replaying actions. The test is still trying to achieve the same outcome, but the product now expects a different sequence to get there.

Typical patterns include:

  • A signup flow that starts requiring a phone number for accounts in certain countries
  • A workspace creation flow that now asks you to choose a plan before you can invite teammates
  • A settings page that moved from a single form to a tabbed interface

Conditional branching and navigation restructures take the longest to fix because they often require updating test data, adding new assertions, or restructuring multiple test files.

Breakdown by flow change type

Flow change type Share of flow failures Median fix time (human)
New required step38%1.5 hours
Removed / merged step21%1.2 hours
Conditional branching24%2.1 hours
Navigation restructure17%2.8 hours

3. Environment instability: 22% of failures

Typical causes include:

  • Network instability or slow external APIs
  • Dependent services down or degraded
  • Bad test data or corrupted seed databases
  • Bad deployments or rollout misconfigurations

Breakdown by environment issue

Test data problems are the leading cause. Shared test environments where multiple CI jobs pollute the database, or seed data that drifts from production schemas, create failures that are technically environment issues but feel like application bugs.

Environment instability breakdown

4. Loading and timing issues: 19% of failures

Loading and timing issues show up as race conditions between UI rendering and test actions, dynamic content that appears only after an API call, and skeleton screens or spinners masking real content.

This is often labeled "flakiness," but in practice most of it is predictable based on DOM and network patterns.

Typical examples:

  • Clicking a button before a React effect attaches the handler
  • Reading table rows before the data fetch completes
  • Asserting on page text while a loading placeholder is still visible

Breakdown by timing pattern

Timing pattern Share of timing failures Median fix time (human)
Element not yet interactive42%25 min
Data not yet loaded31%35 min
Animation / transition interference15%40 min
Async state race conditions12%1.2 hours

Most timing issues can be solved with better wait strategies. The hardest cases involve async state management where the UI renders in an intermediate state that looks correct but isn't.

These timing categories echo a pattern from Checksum's State of AI code report. When engineering leaders traced a rolled-back AI change to its root cause, 21.8% pointed to AI logic errors and 17.9% said the AI couldn't see how the system's pieces interacted.

That's the same failure our flow-change and environment-instability data describes, just measured from the other side: a survey asking engineers why AI code broke in production, versus telemetry showing why automated tests broke in CI. Two different methods both conclude that the hardest failures to prevent are the ones where the system making the change doesn't have full visibility into how the product actually behaves.

Detection to resolution: average times

With failures categorized, we can assess how long it takes from first failure detection to a stable green run.

Human-only resolution times

Flow changes have the longest tail. Complex flow changes (those affecting multiple tests or requiring coordination with product teams to understand intended behavior) can stretch to 5+ hours.

Time breakdown: what engineers actually do

For a typical failure requiring human intervention, investigation dominates. Engineers spend more time figuring out what broke and why than actually fixing it. This is the phase where AI assistance provides the most leverage.

What changes with AI

Checksum handles failures in two stages: a real-time auto-recovery loop that protects CI signal, and a slower auto-healing loop that updates the underlying tests.

Real-time auto-recovery

Real-time auto-recovery runs at the moment a test fails. It replays the scenario interactively, tries small targeted adjustments, and decides whether the product is broken or the test is.

Real-time auto-recovery does not necessarily change the test code. Its job is to get you a reliable green or red and a clear diagnosis. This is effectively manual debugging carried out by an assistant inside the run instead of in an editor.

For customers using it in CI, around 80% of failures are recovered in real time, either the test passes again with lightweight fixes, or the system gathers enough evidence to flag a real product bug. Added wall-clock time is usually measured in tens of seconds, not minutes.

80%
Real-time
auto-recovery in CI

Auto-healing

Once the real-time layer has done its job and the team knows whether they're looking at a bug or a broken test, auto-healing takes over. Auto-healing works on the test suite itself, proposing and applying code changes that keep future runs green.

About 70% of test issues are fixed completely autonomously by auto-healing. For those that require a human in the loop to review and approve suggested patches, around 98% of test issues are fully resolved in under 10 minutes from first failure to stable green run.

70%
Test issues fixed
with auto-healing
98%
Resolved with less than
10 minutes of engineering time

Success rates by root cause

Selector changes are the sweet spot for autonomous repair. The intent is unchanged and the fix is mechanical. Flow changes have a lower autonomous repair rate because they often require judgment about whether the test should adapt to the new flow or whether the flow change itself is a bug.

Root cause Fully autonomous With human review
Selector changes91%99%
Flow changes52%96%
Environment instability48%94%
Loading / timing84%99%

Walking through a flow change test failure

One Checksum customer, a large, publicly traded enterprise software company, hit an issue that illustrates how AI agents approach and repair flow change test failures.

This particular test invited a new user to a shared workspace. But underneath it, the product had changed: the team added a second verification email to the signup flow. The test, which had been built around the original single-email version, grabbed the first email sitting in the inbox (the old one-time code) and failed.

Checksum's real-time layer recognized that something about the test itself no longer matched the product. Auto-healing took over and the agent worked out that two emails now arrive for the same signup, but the correct one is whichever arrives after the profile is submitted, not whichever arrives first. As a legitimate product change, not a bug, the test was rewritten to select the email by timestamp, open the right verification link, and continue. The agent then confirmed the invited user actually reached the workspace before closing the loop.

Telling "this broke" from "this changed on purpose" takes production-level context that an AI coding agent can't get from the code alone, which is why flow changes only clear full autonomy about half the time. This is what it looks like when they do.

Predictable web automation failure is an opportunity

The majority of failures come from selectors and DOM structure—categories where AI can understand the intent and propose fixes with high accuracy. The truly flaky stuff (networks, environments, and bad deployments) accounts for under a quarter of failures in our data set.

This is good news. It means most failures are in the part of the stack that AI can understand and fix: the code that glues your tests to your product.

PART 3

The true cost of maintaining a test suite

Why test maintenance is more expensive than you think, and how to calculate what you're actually spending.

Test automation is sold on efficiency: write the test once, run it forever. But the reality is different. Tests break constantly, and someone has to fix them. That someone is usually your most experienced engineer, and they're not cheap.

We analyzed maintenance costs across hundreds of customer teams to understand what test maintenance actually costs, where the money goes, and how AI-assisted maintenance changes the economics.

The hidden cost problem

Test maintenance is invisible until it isn't. Unlike feature work, it doesn't show up on roadmaps. Unlike incidents, it doesn't trigger alerts. It lives in the space between, a tax on engineering time that accumulates quietly.

When we ask engineering managers how much time their team spends on test maintenance, the typical answer is "not much" or "a few hours a week." But when we instrument the actual time, the numbers are consistently higher.

The gap exists because maintenance is fragmented. Ten minutes here debugging a selector. Twenty minutes there waiting for CI to confirm a fix. An hour lost to a flaky test that turns out to be a real bug. None of these feel significant in isolation, but together they add up.

The cost model

For the purpose of this study, the cost of test maintenance is modelled as:

Monthly maintenance cost = Failures per month × Time per failure × Engineer hourly rate

Failures per month

From our analysis of 1M+ test runs, teams without AI-assisted maintenance see a median of 14.8 failures per 100 test runs (see the section on overall failure rates for more information).

Suite size Tests Daily runs* Failure rate Monthly failures
Small1001514.8%222
Large5002514.8%1,850

*Daily runs refers to business days

Cost and time per failure type

Using an effective engineer rate of $150/hour for US-based teams, we can calculate cost per root cause. Note: a fully loaded experienced engineer is priced at $125–$175/hour.

Not all failures take the same time to fix. When you weigh the findings of our root cause analysis by frequency and add overhead for context switching, CI wait times, and occasional misdiagnosis, the realistic all-in time per failure for human-only maintenance is approximately 1.3 hours.

Flow changes are by far the most expensive to fix manually. They require understanding product intent, often span multiple test files, and may need coordination with product teams. A single flow change that breaks five tests can easily consume a full day of engineering time.

Test automation is sold on efficiency: write the test once, run it forever. But the reality is different. Tests break constantly, and someone has to fix them. That someone is usually your most experienced engineer, and they're not cheap.

Root cause Share Avg fix time Cost per failure
Selector changes32%45 min$113
Flow changes27%2.1 hours$315
Isolated flakiness22%1.5 hours$225
Loading / timing19%40 min$100
Blended average1.3 hours$195

Quantifying human-only maintenance

Monthly costs

Applying the blended $195 per failure reveals that a team with 500 tests is spending the equivalent of several engineers just keeping tests green. They're not writing new tests or improving coverage. That's time spent purely on maintenance.

Suite size Monthly failures Cost per failure Monthly cost Annual cost
Small (100 tests)222$195$43,290$519,480
Large (500 tests)1,850$195$360,750$4,329,000

Monthly engineering hours

Converting to time makes the burden even clearer. In practice, teams with large suites don't fix every failure. They triage aggressively, disable flaky tests, and accept lower coverage. They pay the price in slower releases, reduced confidence, and technical debt.

Suite size Monthly failures Hours per failure Monthly hours FTE equivalent
Small (100 tests)2221.32891.7
Large (500 tests)1,8501.32,40513.9

AI-assisted maintenance with Checksum

Checksum changes the cost equation by handling most repairs autonomously and reducing human involvement to review and approval.

Time per failure with Checksum

With AI-assisted maintenance, the time profile changes dramatically. The weighted average is approximately 5 minutes per failure, which represents a 94% reduction in human time.

Resolution path Share of failures Time required
Fully autonomous (no human needed)70%0 min
Human review only28%10 min
Manual intervention required2%1.3 hours

Failure rate reduction

Beyond faster fixes, Checksum reduces the failure rate itself. From our data, AI-maintained suites see median failure rates of 2.7 per 100 runs versus 14.8 for manual maintenance. This is an 82% reduction.

The transformation happens because auto-healing fixes fragile selectors before they cause repeated failures, better wait strategies reduce timing-related flakiness, and the system learns application-specific patterns over time.

AI-assisted maintenance monthly costs

Across both small and large suites, Checksum's auto-healing cuts failures reaching a human by roughly 82%. The fixes that remain take minutes of review instead of a full manual fix. That combination turns maintenance cost from an hours-per-failure problem into a minutes-per-failure one, driving the total cost down by about 99% regardless of suite size.

Suite size Baseline failures With Checksum Human time Monthly human cost
Small (100 tests)222413.4 hours$510
Large (500 tests)1,85033828 hours$4,200

Side-by-side comparison

Suite size Human-only monthly With Checksum monthly Monthly savings Annual savings
Small (100 tests)$43,290$510$42,780$513,360
Large (500 tests)$360,750$4,200$356,550$4,278,600

Secondary costs: what the model misses

Looking at the direct cost model reveals how much time engineers spend fixing tests. But there are several other costs that are harder-to-quantify.

Blocked releases

When CI is red, deployments wait. In our customer data, average time from test failure to deployment-ready is 3.2 hours for human-only maintenance versus 18 minutes with AI-assisted maintenance. Teams without AI maintenance report an average of 4.2 blocked or delayed releases per month.

Context switching

Engineers pulled into test maintenance lose flow state. UC Irvine research suggests that it takes 23 minutes to return to a task at the same level of focus after a context switch. For a team handling 10 failures per day, that's almost four hours of daily lost productive time.

Trust erosion

Flaky tests that cry wolf train engineers to ignore failures. Once trust erodes, real bugs slip through because failures are assumed to be test issues, engineers stop writing tests for new features, and coverage plateaus or declines over time.

This isn't unique to test suites. Checksum's State of AI code report found 45.7% of engineering leaders saw AI-code issues drive up support load over the past year. Once a system's signal can't be trusted, the burden doesn't disappear. It moves downstream to whoever has to clean up after it.

Opportunity cost

Every hour spent on test maintenance is an hour not spent on new feature development, performance optimization, security improvements, or technical debt reduction. That trade-off shows up in Checksum's State of AI code report too: 42.9% of engineering leaders reported losing engineering hours to rework in the last 12 months because of AI-generated code issues. It's the same hours-per-failure math seen here but from a different angle: time spent maintaining what should already work is time that engineers can't spend building something new.

How to track your test maintenance cost

To understand and control test maintenance costs, it's vital to measure failure, cost, and health metrics. Many teams track none of these, and the ones that do are consistently surprised by what they find.

Failure metrics

  • Failures per 100 test runs (overall and by root cause)
  • Time from failure detection to green CI
  • Repeat failure rate (same test failing multiple times before stable fix)

Cost metrics

  • Engineer hours spent on test maintenance
  • Deployment delays attributable to test failures
  • Test-disable rate (tests turned off due to flakiness)

Health metrics

  • Test coverage trend over time
  • Ratio of new tests written to tests disabled
  • Mean time to diagnose (test bug vs product bug)
CASE STUDY

A case study in test maintenance cost

The customer referenced in Part 2 is a working example of this math. Over a 30 day period, Checksum diagnosed, repaired, and opened a pull request for 156 broken tests in that company's suite. Calculating based on this report's blended manual cost of $195 per failure, that's the equivalent of more than $30,000 in engineering time avoided in a single month for one customer's suite.

$30K
Money saved on test
maintenance every month

The volume of repaired tests isn't the whole story. Over the same period, Checksum logged 41 product bugs for this customer, each with its root cause written directly into the test. Those tests remain red until the team ships a real fix, because turning a test green isn't the goal. Knowing which tests should stay red is.

CONCLUSION

Changing the economics of test maintenance

Test maintenance is more expensive than most teams realize. The cost compounds with test suite size, and scales faster than linear because larger suites have more interdependencies and more complex failure modes.

By handling 70% of repairs autonomously and reducing the rest to quick reviews, AI-assisted maintenance tools like Checksum compress multi-hour debugging sessions into minutes. The direct savings are significant. The indirect savings—faster releases, fewer interruptions, sustained trust in automation—are larger still.

This report and our companion State of AI code report measure the same shift from two different angles. One tracks what breaks and what it costs to fix. The other tracks how confidence in AI-written code is rising faster than teams' ability to verify it. Together, they point to the same idea: CI/CD automated how code gets delivered. Continuous verification automates how it gets proven.

The question isn't whether you can afford AI maintenance. Given the numbers, the question is whether you can afford not to have it.

Where This Goes

Verify AI code before it's reviewed, not after it ships

Checksum's E2E Agent automatically generates, runs, and auto-heals a production-like test suite, so AI code is verified before anyone reviews it, not debugged after it ships.

See how it works