It’s been 25 years since we started talking about Continuous Integration, 15 years since we started talking about Continuous Deployment. They’re so standard now we don’t even say them, CI/CD just rolls off the tongue.
Today I’d like to argue for adoption of a third continuous process: Continuous Verification

What CI/CD was built for, and where it stops
CI was designed around a simple guarantee: if your tests pass (not just on your local machine), your code is mergeable. Fast, scoped to the diff, runs on every PR. That constraint is a feature, not a limit.
CD automated the handoff from "tests passed" to "code is live." It made deployment boring in the best sense. Predictable. Repeatable. Less scary.
Neither was designed to monitor whether the running system holds up over time. Whether flows that worked last week still work today. Whether a dependency change three deploys ago quietly broke something nobody's tested since.
That's a different problem. It needs a different name.
The third stage already exists. It just doesn't have a fitting name yet
Integration, Deployment, Verification. Three action nouns, same grammatical shape, sequential stages in the pipeline. ‘CV’is the natural third member of the family. -A recognition of something that's already happening (or should be) in every production system.
"Is it still working? Is my customer my alert system?"
“The beginning of wisdom is to call things by their proper name.” - Confucius
CV runs continuously against the live system. Where CI tests the diff, CV tests the whole application. Every critical flow, every interaction path, every user journey that matters. More than just smoke tests
The bugs that slip through CI aren't random. They're the ones that require state. A user who signed up three days ago trying a flow that changed yesterday. An API endpoint that returns the right response in isolation but breaks when chained with two others.
CI wasn't built to catch those. It wasn't supposed to be. CV is.
Why now
A year ago, the typical engineering team was shipping maybe 20-30 PRs a week per developer. That was already a lot to keep quality intact across.
Now, with AI coding agents in the loop, the same developer ships 50, 80, 100. Code volume has gone up several times over. QA headcount hasn't.
Manual testing doesn't keep up with AI output. Outsourcing buys you a linear cost reduction against an exponentially growing problem. You're paying humans to do a job whose volume is compounding.
Generating tests with Cursor or Claude Code helps at the margins, but "help me write a test for this function" is on-demand and human-driven. The moment you close the tab, the coverage stops.
What actually scales is verification that runs in the background. On every PR, on every deploy, and at rest too. Tests that heal themselves when the application changes, the way your dream human QA used to.
Not a new test tool. More like hiring a team.
How Checksum fits
On every PR, the CI agent generates and runs targeted tests for what changed. The E2E agent maintains a live graph of your application's flows and keeps coverage green as the UI evolves, healing selectors and test logic on its own.
70% of test failures get resolved without anyone on the customer's team touching them. Teams used to spend 90 minutes triaging a single test failure. Now they don't.
Tests ship as Playwright code your team owns. If you drop us tomorrow, they still run, they’re still your code. The whole thing is wired into Claude code + friends with the /checksum slash commands, verification can happen inside the same workflow where the code gets written - but not the same agent grading it’s own homework
Our Goal is to kill the Friday afternoon panic when a 30k-line AI authored PR lands.

