AI writes the code. Someone still has to make sure it works.
That "someone" has been the bottleneck for the last two years. Senior engineers spending their days reviewing AI-generated PRs, updating test selectors, triaging failures, and trying to tell the difference between a real bug and a broken test. The faster the coding agents get, the wider the gap grows.
Today we're launching the Continuous Quality Agent: the verification layer for AI-generated code. It detects gaps in coverage, generates tests for specific flows, runs them against your live application, and heals broken tests autonomously. Fine-tuned on more than 1.5 million test runs. Every test is standard Playwright code, committed to your own repo as a pull request. No proprietary format. No lock-in.
70% of test failures resolve without an engineer touching them.
The best way to understand what the agent does is to watch it work. Here are four short demos covering the full loop.
1. Detection: finding the coverage you don't have
The first job of the agent is to find the gaps in test coverage.
In this demo, we're showing how to detect tests for a new feature. The setup takes about ten seconds: create a folder, click "Detect Test Cases," mode, and kick off the session.
Then walk away.
The session runs hands-off in the background. When it's done, there are 37 test cases waiting in a pull request, each with a detailed explanation of what will be automated, what data setup is needed, what cleanup looks like, and every step the test should cover. The agent figures out the depth and shape of each test based on the feature itself.
Detection is the part most teams skip because it's tedious and easy to deprioritize. The agent doesn't skip it.
2. Generation: from test plans to Playwright code
Once the agent detects coverage gaps, it generates Playwright tests and submits them to your repo as a pull request.
Select the files, click "Generate," and the agent starts cloning the repo and writing tests. In the video above 33 tests were generated. 31 passed. 2 failed because they caught real bugs.
Every test file lands in TypeScript, with data setup, cleanup, and full isolation built in. The output is standard Playwright code, dropped into a pull request you review like any other diff. You own the tests. You can read them, edit them, run them anywhere. If you ever decide to stop using Checksum, your test suite goes with you.
This is the deliberate design choice that matters most. The industry is full of "AI testing platforms" that lock your tests inside a proprietary system you can't take with you. We made the opposite call.
3. Auto-healing: when tests break, the agent fixes them
This is where most teams feel the pain.
A product change ships, selectors shift, half the test suite goes red, and someone has to spend their morning figuring out which failures are real bugs and which are just maintenance debt. Multiply that by every sprint.
In this demo, a test run completes with six failures. The agent moves straight into triage, analyzes what happened in each failing test, and surfaces the root cause in a "recovered" state. That alone tells the team why each test failed.
Then comes the second phase: writing the actual code fix. The agent opens a pull request with the updated tests, including a fix for a test that was tagged as a known bug. You review the diff, approve it, merge it. Test suite back to green.
This is the "70%" we talk about. Most test failures don't need a person. They need an agent that knows the test, the product, and how to reconcile them.
4. Healing inside CI/CD: feedback on every PR
The fourth video is the one that changes how teams actually use this in practice.
A developer commits a change, pushes the branch, opens a PR. The action triggers healing automatically. The runner kicks off the testing suite, the agent wakes up if anything breaks, and a comment is posted directly on the PR with a real-time status.
Six failures, one pass. Triage in progress. Status moves from running to recovered, with annotations on every failure explaining the root cause. Then the agent updates the code and submits the fix as a separate PR.
The whole loop runs inside the developer's existing workflow. No dashboard to check. No engineer waiting around. Just the right comment on the right PR.
What this changes
For most engineering teams, end-to-end testing is a thing you put off until it hurts, then a thing you do badly because the maintenance cost eats every improvement you make. The Continuous Quality Agent is what closes that gap.
You ship more code. The agent keeps the tests current. Failures get triaged and fixed at machine speed. Engineers spend their time building, not babysitting.
Or, as Ron Alexssen, Engineering Manager at Counterpart, put it: "For less than half the salary cost of an offshore developer, I have the impact of a full QA team."
That's the difference between AI as a copilot and AI as infrastructure.
See the agent on your own codebase
The Continuous Quality Agent is available now. Book a demo and we'll show you what it looks like running against your application, your tests, your CI.
