Skip to main content
Developer reference
IDs, outcomes, and verdict values used by the API, CLI, and MCP

IDs you’ll meet

Most integration mistakes come from passing the wrong ID. This table covers each one.
Healing needs the run UUIDTo heal a run that already finished, pass testRunId set to the run’s UUID: the runId returned when you started it, or the id from GET /test-runs/latest. The job name won’t work.

Outcomes and verdicts (API values)

The enum values your code will read from run, result, and report payloads.

Test outcomes

These terms are used consistently across the web app, CLI, API, and notifications:

Run verdict

For CI gating, the API computes one verdict per run: pass, fail, or pending. Gate on verdict once isTerminal is true, not on status or on the raw counts. An empty test selection (executedCount: 0) is a fail, and a sharded run only gets a verdict after its shards merge. See Run status.

Report verdicts

Separately, people and automation can submit a per-test verdict on a run report: bug, recovered, healing, or triage, each with an annotation. See Submit verdicts.
i
Glossary
The objects you work with in Checksum

How the pieces fit

ProjectAPI key, environments, Git
→
CollectionFeature area
→
Test flowOne user journey
→
Story + test.checksum.md + .spec.ts
→
Test runResults, artifacts, verdict
Agent sessions are the work units that detect flows, generate tests, and heal them. Bug entities group the failures that sessions classify as real application bugs.

Project and configuration

Project

Everything in Checksum belongs to a project, which usually means one application under test. A project has exactly one API key, one or more environments, a connected tests repository, and a connected code repository. Checksum creates your project during onboarding.

Environment & test users

An environment is a deployment of your app to test against (e.g. UAT, staging), with an environment URL and an optional login URL. Each environment has one or more test users, the credentials the agent and test runs use to log in, often one per role (admin, viewer). See Environments & Test Users.

Tests repository & code repository

The tests repository is where your Playwright tests live. Checksum writes to it only by opening pull requests. The code repository is your application source, which Checksum only reads, to improve detection and generation. They can be the same repo. See Test Repository & Config.

Repo mirror

Checksum keeps a synchronized view of your repositories through webhooks (pushes, PR events, app installation changes). It reads your existing tests so it doesn’t generate duplicates, and it writes back only through PRs. Your repository is always the source of truth. Configuration works differently: settings in the web app don’t sync to your repository. Your tests read from a .env file that Checksum maintains and you download with npx checksumai dotenv --download (see Environments). More on the repo mirror →

Tests

Collection

A group of related test flows, such as “Checkout”, “User Management”, or “Settings”. Collections organize the suite by feature area. You can run a whole collection by ID from the API (POST /execution/collection/:id) and filter dashboards and bug lists by collection.

Test flow (user story)

One test scenario inside a collection, describing a user journey like “User can create an account” or “Admin can export a report”. A flow has a title, a description of its steps, and a start URL. Flows come from detection or are created manually, and they are what Checksum generates tests for. You can also skip writing flows and record a walkthrough with the Chrome extension, which the agent uses as context to generate one test or a batch of tests, one per flow.

Story file & test file

Each generated test is two files: a story (.checksum.md, a human-readable spec with frontmatter, data setup and cleanup, and steps) and a test (.checksum.spec.ts, Playwright code using Checksum fixtures). See Story & Test Format.

Tags

Labels on tests (e.g. smoke, checkout) that appear in results and can filter dashboards, bug lists, and notification reports. A grep such as @smoke is a common way to select a subset for a run.

Agents

Agent session

A running instance of the AI agent doing one piece of work: detecting flows, generating a test, or healing failures. Sessions move through a lifecycle (Initializing → Cloning → Running → Completed/Failed) and can pause for your input or approval. See Agent Sessions.

Standard vs Deep mode

Every pipeline (detection, generation, and healing) runs in Standard mode (fast and fully autonomous; the default) or Deep mode (interview and plan first, then a knowledge-base update and more thorough review). See Deep vs Standard Modes.

Batch

API-triggered generation and healing return a batchId. A batch groups one or more sessions started by the same request. For example, one heal request can fan out into a session per failing test. Poll the batch until allTerminal is true.

Runs and outcomes

Test run

One execution of some or all tests, via the CLI, CI, the GitHub Action, or the REST API. A run records per-test outcomes, branch and commit, and artifacts (videos, screenshots, HAR files, Playwright traces), and appears under Test Results. A run can be sharded across parallel machines and merged into one result (see Sharding).

Bug entity

A trackable bug (e.g. BUG-42) that groups every failing test caused by the same problem, so you triage, comment on, and resolve it once. Statuses: Needs Triage, Confirmed, Fixed, Not Bug, Snoozed. See Feature Health Dashboard.

Test health

A rolled-up status based on recent run history, not just the latest run: Passing or Failing, with per-test recent-run history to show intermittent failures. Result payloads also carry a healthStatus such as healthy or flaky.

API Keys & Authentication

Base URL, versions, and the typical API flow.

How Test Generation Works

From flows to PRs.

Results, Reports & Traces

Where outcomes show up.