Developer reference
IDs, outcomes, and verdict values used by the API, CLI, and MCP
IDs you’ll meet
Most integration mistakes come from passing the wrong ID. This table covers each one.Outcomes and verdicts (API values)
The enum values your code will read from run, result, and report payloads.Test outcomes
These terms are used consistently across the web app, CLI, API, and notifications:Run verdict
For CI gating, the API computes oneverdict per run: pass, fail, or pending. Gate on verdict once isTerminal is true, not on status or on the raw counts. An empty test selection (executedCount: 0) is a fail, and a sharded run only gets a verdict after its shards merge. See Run status.
Report verdicts
Separately, people and automation can submit a per-test verdict on a run report:bug, recovered, healing, or triage, each with an annotation. See Submit verdicts.
Glossary
The objects you work with in Checksum
How the pieces fit
ProjectAPI key, environments, Git
→
CollectionFeature area
→
Test flowOne user journey
→
Story + test
.checksum.md + .spec.ts→
Test runResults, artifacts, verdict
Project and configuration
Project
Everything in Checksum belongs to a project, which usually means one application under test. A project has exactly one API key, one or more environments, a connected tests repository, and a connected code repository. Checksum creates your project during onboarding.Environment & test users
An environment is a deployment of your app to test against (e.g. UAT, staging), with an environment URL and an optional login URL. Each environment has one or more test users, the credentials the agent and test runs use to log in, often one per role (admin, viewer). See Environments & Test Users.Tests repository & code repository
The tests repository is where your Playwright tests live. Checksum writes to it only by opening pull requests. The code repository is your application source, which Checksum only reads, to improve detection and generation. They can be the same repo. See Test Repository & Config.Repo mirror
Checksum keeps a synchronized view of your repositories through webhooks (pushes, PR events, app installation changes). It reads your existing tests so it doesn’t generate duplicates, and it writes back only through PRs. Your repository is always the source of truth. Configuration works differently: settings in the web app don’t sync to your repository. Your tests read from a.env file that Checksum maintains and you download with npx checksumai dotenv --download (see Environments). More on the repo mirror →
Tests
Collection
A group of related test flows, such as “Checkout”, “User Management”, or “Settings”. Collections organize the suite by feature area. You can run a whole collection by ID from the API (POST /execution/collection/:id) and filter dashboards and bug lists by collection.
Test flow (user story)
One test scenario inside a collection, describing a user journey like “User can create an account” or “Admin can export a report”. A flow has a title, a description of its steps, and a start URL. Flows come from detection or are created manually, and they are what Checksum generates tests for. You can also skip writing flows and record a walkthrough with the Chrome extension, which the agent uses as context to generate one test or a batch of tests, one per flow.Story file & test file
Each generated test is two files: a story (.checksum.md, a human-readable spec with frontmatter, data setup and cleanup, and steps) and a test (.checksum.spec.ts, Playwright code using Checksum fixtures). See Story & Test Format.
Tags
Labels on tests (e.g.smoke, checkout) that appear in results and can filter dashboards, bug lists, and notification reports. A grep such as @smoke is a common way to select a subset for a run.
Agents
Agent session
A running instance of the AI agent doing one piece of work: detecting flows, generating a test, or healing failures. Sessions move through a lifecycle (Initializing → Cloning → Running → Completed/Failed) and can pause for your input or approval. See Agent Sessions.Standard vs Deep mode
Every pipeline (detection, generation, and healing) runs in Standard mode (fast and fully autonomous; the default) or Deep mode (interview and plan first, then a knowledge-base update and more thorough review). See Deep vs Standard Modes.Batch
API-triggered generation and healing return abatchId. A batch groups one or more sessions started by the same request. For example, one heal request can fan out into a session per failing test. Poll the batch until allTerminal is true.
Runs and outcomes
Test run
One execution of some or all tests, via the CLI, CI, the GitHub Action, or the REST API. A run records per-test outcomes, branch and commit, and artifacts (videos, screenshots, HAR files, Playwright traces), and appears under Test Results. A run can be sharded across parallel machines and merged into one result (see Sharding).Bug entity
A trackable bug (e.g.BUG-42) that groups every failing test caused by the same problem, so you triage, comment on, and resolve it once. Statuses: Needs Triage, Confirmed, Fixed, Not Bug, Snoozed. See Feature Health Dashboard.
Test health
A rolled-up status based on recent run history, not just the latest run: Passing or Failing, with per-test recent-run history to show intermittent failures. Result payloads also carry ahealthStatus such as healthy or flaky.
Related
API Keys & Authentication
Base URL, versions, and the typical API flow.
How Test Generation Works
From flows to PRs.
Results, Reports & Traces
Where outcomes show up.