At a glance
Reference for AI: running tests at a glance
Reference for AI: running tests at a glance
- Base URL:
https://api.checksum.ai/public-api/v1/. Only grep (and its job-name status) useshttps://api.checksum.ai/public-api/v2/. - Auth: every REST call sends
Authorization: Bearer $CHECKSUM_API_KEY. Key location: Settings → Project Settings. - Gating CI: poll
GET /execution/status/run/{runId}untilisTerminalistrue, then pass only whenverdictis"pass". - Heal on failure: add an
autoHealobject to any execution request, or pass--cksm-auto-healto the CLI. - No cancel: API-triggered runs can’t be cancelled through the public API.
- Sharding:
shardCount2–40. Requireschecksumai4.4.0+ on the tests branch.
Before you run the CLI
Checksum sets up your tests repository during onboarding. To run it with the CLI, clone the tests repo and prepare it once:dotenv --download writes a .env file with your project’s environment settings: environment URL, login URL, credentials, and custom variables. Your API key is in Settings → Project Settings (see API Keys). Checksum maintains the file’s values, so to change them, contact your Checksum team and then download it again. Any automation that starts a Checksum test run or AI generation should download the latest .env first (see Environments).
To check your setup, run the built-in example test. It’s a quick check that your login works:
Run tests with the CLI
The Checksum CLI (checksumai) runs your generated Playwright tests. It downloads your environment configuration, runs the tests, uploads the results to the dashboard, and can hand failures to auto-healing. It ships in the @checksum-ai/runtime npm package, and every command runs with npx checksumai. Because it wraps Playwright, your tests also run under plain Playwright.
Everyday commands
CI=true), results upload to the dashboard automatically. See What gets reported.
Useful flags
Most runs need no flags at all. The ones you’ll reach for:-g "pattern"runs only the tests whose name matches.--cksm-auto-healsends failures to auto-healing when the run ends. In GitHub Actions and GitLab CI, Checksum detects the repository, branch, and pull request on its own, so this flag alone is usually enough.--cksm-affectedruns only the tests a change affects (below).--cksm-rerun-failedre-runs only what didn’t pass last time (below).
Run only the tests a change affects
For pull-request checks, you usually don’t need the whole suite.--cksm-affected looks at which files changed since a base branch, asks Checksum which tests those files affect, and runs just those:
fetch-depth: 0 in actions/checkout). It can’t be combined with -g or --cksm-rerun-failed.
Re-run only what failed
After a run with failures, you can retry just the test files that didn’t pass (failed or recovered), without running everything again:Update the CLI
Installing your tests repository’s dependencies installs the CLI (see Before you run the CLI). To update to the latest release, which you need before using sharding (minimum4.4.0), run this and commit the result:
Other CLI commands
Reference for AI: Checksum CLI (npx checksumai)
Reference for AI: Checksum CLI (npx checksumai)
checksumai (npm dependency of the tests repository). Invocation: npx checksumai <command>. Wraps Playwright; generated tests also run under plain Playwright. Update: npm install checksumai@latest, then commit package.json / lockfile to the tests branch; sharding requires 4.4.0+.Setup per checkout: npm install, npx playwright install --with-deps, npx checksumai dotenv --download --api-key=$CHECKSUM_API_KEY. Verify with npx checksumai test -g "example" (checks login). Re-download after Checksum updates the variables, and in any automation that starts a Checksum test run or AI generation.Commands
npx checksumai test flags
--cksm-auto-heal alone is usually enough.--cksm-affected behavior and constraints
Three steps: (1) compute changed files between the current checkout and the base ref; (2) call POST /public-api/v1/affected-tests; (3) run Playwright with an internal --grep over the returned test IDs.--cksm-rerun-failed behavior
Re-runs only test files that didn’t pass (failed or recovered). Without an ID, fetches per-test results via GET /public-api/v1/test-runs/latest; with =$TEST_RUN_ID, uses that run. Resolves failing file paths and passes them to Playwright. If every test passed, exits successfully with nothing to run. Can’t be combined with --cksm-affected or -g.Reporting: with hostReports enabled (default true when CI=true), results upload to the dashboard.Run from GitHub Actions
On GitHub, the simplest option is the Checksum GitHub Action. It starts a cloud run from your workflow and calls the REST API for you, so you don’t need Playwright on your runner. Store your API key as the repository secretCHECKSUM_API_KEY:
Reference for AI: GitHub Action summary
Reference for AI: GitHub Action summary
checksum-ai/test-run-action@v2. Calls the REST execution endpoints. Required input: api-key: ${{ secrets.CHECKSUM_API_KEY }}. Selection modes: grep, affected, suite-ids, test-ids, collection-id. Other inputs: wait, sharding, env overrides, auto-heal. Full input table: CI/CD Integration → GitHub Action.Start a cloud run from the REST API
With the REST API, Checksum runs your tests in its own cloud, which works from any CI system or script. Each call starts a run and immediately returns arunId. You then use that ID to check the result. All calls send Authorization: Bearer $CHECKSUM_API_KEY (see Base URL & versions). The v1 execution endpoints run against the project’s configured branch and environment.
autoHeal, which heals failures automatically when the run ends (Auto-Healing), and shardCount, which splits the run across parallel machines (Sharding).
Run the whole suite
acme-co/my-tests.
Run a collection
Run specific tests
Run tests that match a name or tag
The most flexible option, and the one to use for pull-request checks. Besides a name or tag pattern such as@smoke, it can check out a specific branch of your tests repository and point the run at a preview deployment:
BASE_URL) can be overridden. Names reserved by Checksum are rejected. An older v1 version of this endpoint accepts only the pattern. Use v2 for new integrations.
Find the tests a change affects
Send the list of files a change touched (for example, the output ofgit diff --name-only) and Checksum returns the tests most likely affected. Then run just those with Run specific tests. The CLI’s --cksm-affected flag does both steps for you.
Reference for AI: execution and affected-tests endpoints
Reference for AI: execution and affected-tests endpoints
https://api.checksum.ai/public-api/v1/ (all execution endpoints except grep) and https://api.checksum.ai/public-api/v2/ (grep). Headers on every call: Authorization: Bearer $CHECKSUM_API_KEY, Content-Type: application/json. v1 execution endpoints run against the project’s configured branch and environment.Execution response (all execution endpoints)
Optional fields accepted by every execution endpoint
POST https://api.checksum.ai/public-api/v1/execution/suite
Runs the full test suite in Checksum’s cloud. Body optional: autoHeal, shardCount. Example body:{ "runId": "9f2c7a4e-8b31-4d6a-a2f0-3c5e1b7d9a42", "name": "job-name-12345", "sharded": false }. Next: GET /public-api/v1/execution/status/run/{runId}.POST https://api.checksum.ai/public-api/v1/execution/collection/{id}
autoHeal, shardCount. Response: execution response. Next: GET /execution/status/run/{runId}.POST https://api.checksum.ai/public-api/v1/execution/tests
GET /execution/status/run/{runId}.POST https://api.checksum.ai/public-api/v2/execution/grep
Runs every test whose name matches a pattern. The only execution endpoint that can check out a specific branch and inject per-run environment variables; use it for PR checks against preview deployments.POST https://api.checksum.ai/public-api/v1/execution/grep still exists with a smaller body (grep only). New integrations should use v2.Response: execution response. Next: poll GET /public-api/v1/execution/status/run/{runId}; pass the CI check only when verdict is "pass".POST https://api.checksum.ai/public-api/v1/affected-tests
Returns the Checksum test IDs most likely affected by a set of changed source files.affectedTestIds as testIds to POST /public-api/v1/execution/tests. The CLI flag --cksm-affected does both steps.Check whether a run passed
A cloud run takes a few minutes. To find out how it went, ask for its status using therunId you got back when you started it. Keep asking until isTerminal is true, then look at verdict: "pass" means the run passed.
Get a run’s status and verdict
Older integrations: status by job name
Before run IDs, integrations checked status by the jobname returned when a run started. That still works for non-sharded runs, but use the run ID for anything new, and always for sharded runs.
Reference for AI: run status endpoints
Reference for AI: run status endpoints
GET https://api.checksum.ai/public-api/v1/execution/status/run/{runId}
Server-computed status of a run by runId from any execution endpoint. Works for sharded and non-sharded runs. Use for all new integrations. Header: Authorization: Bearer $CHECKSUM_API_KEY.isTerminal is true, then pass only when verdict is "pass". Don’t gate on status or on counts. verdict is computed only after sharded results merge; an empty selection is a failure.Next: GET /public-api/v1/test-runs/{id}/results for per-test details, or POST /public-api/v1/auto-heal to heal failures.GET https://api.checksum.ai/public-api/v1/execution/status/{jobName} (legacy)
Also: GET https://api.checksum.ai/public-api/v2/execution/status/{jobName} for v2 grep runs. Non-sharded runs only. Prefer GET /execution/status/run/{runId} for new integrations and all sharded runs. The v2 response adds testRunId (UUID) once terminal; use that value, not jobName, for POST /auto-heal.cancel-in-progress cancels the workflow, not the Checksum run.Example: run the suite, wait for the verdict, heal failures
This script puts the pieces together. It starts a full suite run with auto-heal on failure, then waits for it to finish. If the run fails, healing starts automatically, with no separate heal call. It needscurl, jq, and CHECKSUM_API_KEY in the environment. Set TESTS_REPO and TESTS_BRANCH to your tests repository and its integration branch.
Example: the full script
Example: the full script
autoHeal, call POST /public-api/v1/auto-heal with testRunId set to the run’s UUID (the runId, or the id from GET /test-runs/latest). Don’t use the dispatch name.
Pick which tests to run
Every interface can narrow a run down. Here’s how the same choice looks in each:Environment overrides per run
Runs use the environment settings from your project (Environments → Environment variables). You can override any variable for a single run:.env file.
Run modes and runtime recovery (runMode)
runMode in checksum.config.ts controls how the CLI reacts to a failing step:
useChecksumSelectors (smart selector recovery, default true) and useChecksumAI ({ actions: true, assertions: false } by default, so assertion failures aren’t auto-recovered). See Auto-Maintenance & Recovery and checksum.config.ts.
What gets reported (hostReports)
When hostReports is enabled (default true when CI=true, which most CI providers set automatically), the CLI uploads the following after each run. Local runs keep reports on your machine unless you turn hostReports on.
- Test results: passed / failed / recovered / bug status per test
- Videos of each test’s browser session
- Screenshots at key points and on failure
- HAR files of network traffic
- Playwright traces for step-by-step debugging
Ways to run
Scheduled runs
Periodic runs execute your tests on a recurring, cron-based schedule. You don’t need to trigger runs by hand or rely only on CI events.- Schedule runs in your own CI (
schedule: cronin GitHub Actions,$CI_PIPELINE_SOURCE == "schedule"in GitLab). See CI/CD Integration. - Ask your Checksum team to set up scheduled runs, and to turn on auto-healing after scheduled runs (a project-level setting).
- Create scheduled runs with cron patterns (e.g. daily at midnight, every 6 hours)
- Branch selection: choose which branch to run against
- Test filtering with grep patterns
- Activate / deactivate schedules without deleting them
- Run now: trigger an immediate run outside the schedule
- View results from scheduled runs alongside manual runs
- Auto-heal on failure (project-level setting)
Troubleshooting
The status is passed but verdict is fail
The status is passed but verdict is fail
executedCount. An empty selection (e.g. a grep that matches nothing) is treated as a failure. Fix the pattern or IDs.A sharded run never reaches a final verdict
A sharded run never reaches a final verdict
checksumai older than 4.4.0. Update it, commit, and re-run. Always set a timeout on your polling loop.I cancelled the CI job but the run kept going
I cancelled the CI job but the run kept going
400 on envOverrides
400 on envOverrides
CI or starting with CHECKSUM_. Also, envOverrides is only accepted on v2 grep.--cksm-affected fails in CI
--cksm-affected fails in CI
fetch-depth: 0 in actions/checkout) so the merge base with the target branch is available.The example test fails locally
The example test fails locally
.env, confirm the environment URL and test user in Environments, and make sure your network can reach the environment.