Blog

How We Made AI Test Generation 9× Faster and 6.7× Cheaper

Gal Vered
Date Published:
October 1, 2026

‍

Key takeaways

  1. Architecture built around a model's limits becomes a liability once those limits disappear. Our multi-agent system existed to work around Opus 4.5's 200K context window and weak compaction, and it outlived both.
  2. Splitting work across agents to save context has a cost. Each handoff document carries only part of what the previous agent knew, and every review-and-fix loop adds time.
  3. Newer models with 1M-token context windows can explore an app, write a test, and repair it in a single agent without losing track of the original request.
  4. Give agents only the work that needs judgment. Running a test follows a fixed procedure, so our Python service handles it and the agent focuses on writing and repairing tests.
  5. On our internal benchmark, the new generator was 9× faster and 6.7× cheaper per test.

Why we built test generation around multiple agents

The first iteration of the CQ generation system was during the reign of Claude Opus 4.5. At the time, this was the state of the art in the LLM world and our team built our test generation system around it. While this meant harnessing the thinking and computational power of the model, it also meant building around its limitations. While this was sensible during the time, technology moves incredibly fast, and in late 2026 much of what we built to support Opus 4.5 hinders the effectiveness of new models.

One of the core limitations of Opus 4.5 was its 200,000-token context window. Unlike modern models where context windows sit around 1 million tokens and are much better at long running tasks, the Opus 4.5 window’s small size meant that the requisite application context for the generation of any given test would usually exceed the model’s effective context threshold. At the time, compaction was also not as good as today and models quickly lost orientation and forgot key instructions after a few turns.

How the multi-agent handoff system worked

As a result, our generation system was built using handoffs between multiple agents. For example, an implementation agent, using the Playwright CLI, would browse an application and write a test. At this point, the context of the agent would be near its limits, resulting in degraded responses for any additional prompts. To compensate, a handoff document along with the test itself was passed to another agent for review. That review agent would then hand off suggested fixes to another agent to implement the changes. This cycle would continue until generation was complete.

How the new single-agent generator works

With newer models, however, context windows, compaction and the ability to work on long running tasks while staying true to the original request have increased dramatically. Today agents can hold more information before degrading so we knew it was time to upgrade by tearing the old system down.

Our rebuilt generation system revolves around a single agent with a 1,050,000-token context window and dedicated tools that allow the agent to explore the page, interact with elements and generate accurate tests. 

Once the agent generates the test, our Python backend deterministically stamps it with the required Checksum identifiers and launches the test runner in an isolated environment. If the test passes, the generator’s job is complete. If it fails, the error trace is passed back to the generation agent where the agent will use a set of analysis skills to determine if the test failed because of an application failure or because the test itself was faulty. If the test was faulty, the agent repairs it before passing it back to our Python service to be run again.

Why the agent doesn't run its own tests

It is worth noting that while the agent has enough context to run the test itself, the process to run a test is strictly defined. Asking the agent to initiate it adds model work without improving that procedure.

At best, the agent follows the same procedure at the cost of additional tokens and time. At worst, it chooses the wrong command or configuration, causing an avoidable failure. Our belief is that agents should only be tasked with non-deterministic work and therefore anything that is deterministic, like running a test, should be done via a service. However, “theory only gets you so far,” and to verify the impact of our refactor, we needed to test it.

Results: 9× faster and 6.7× cheaper

The old and new generation systems were tasked with generating a series of tests of varying difficulty on our internal benchmarking application. This web application includes multiple common paths and failures that we tend to find with customers, such as failing form submissions and CRUD operations. What we found was that, on average, our new generator was 9.0 times faster and 6.7 times cheaper than our old system. 

What we'd tell other teams building with agents

The main takeaway from this is to review your product designs, including what tasks are assigned to your agents. With the advent of AI, teams are moving quicker than ever, but with it, less attention is given to pre-existing systems. When those systems rely on technology that is rapidly changing, it is important to ensure that the original design decisions still hold up and to change them if they do not. This includes moving tasks that follow fixed procedures into deterministic code and narrowing agents’ scope to tasks that require judgement.

Gal Vered

Gal Vered is a Co-Founder and CEO at Checksum where they use AI to generate end-to-end Playwright tests, so that dev teams know that their product is thoroughly tested and shipped bug free, without the need to manually write or maintain tests.