Skip to main content

At a glance

  • Secret: store the project API key as a CI secret named CHECKSUM_API_KEY. Checksum sets up the tests repository during onboarding, so there is nothing to initialize.
  • Cross-repo: if the tests live in a different repository than the pipeline, you also need a Personal Access Token with read access to the tests repo.
  • Gate on verdict: pass the build only when verdict is "pass". Don’t gate on the passed/failed counts.
  • No cancel: cancelling a CI job only stops polling. An API-triggered Checksum run continues until it finishes.
  • Sharding requires checksumai 4.4.0 or later on the tests branch.
Developer guide
GitHub Action, the CLI in GitHub Actions and GitLab, the REST API, and auto-heal

Run Checksum with the GitHub Action

The checksum-ai/test-run-action action starts a Checksum run in a single step, with auto-heal on failure built in. The tests run in Checksum’s cloud, so your job doesn’t need Node, Playwright, or browsers. Add your API key as a repository secret named CHECKSUM_API_KEY (from Settings → Project Settings in the Checksum web app; in GitHub, go to Settings → Secrets and variables → Actions), then add this workflow:
.github/workflows/checksum.yml
That’s the whole step. The action finds the source PR on its own (from the event payload on pull_request, or via the GitHub API on push) and passes it to auto-heal, so healing progress shows up as a comment on the right PR. Checksum posts these PR comments itself through its GitHub App, so your workflow doesn’t need its own comment steps or pull-requests: write.
Have a coding agent build itThe CI Setup Prompts are fill-in-the-blanks prompts that have Claude Code, Cursor, or Copilot write this workflow for your repository, including a preview URL, sharding, and auto-heal.
Choose which tests to run with exactly one of these inputs: grep (a name pattern), affected (only tests affected by the PR’s changes), suite-ids, test-ids, or collection-id. The endpoints behind them are described in Running Tests → Execution endpoints, and the complete inputs/outputs reference and changelog are in the action’s README.
@v2 is a breaking releaseAn empty test selection (for example, a grep that matches nothing) now fails the step instead of passing. Read the v2.0.0 release notes for the full list of behavior changes before upgrading from @v1.

Heal failures, with or without a PR

With auto-heal: true, a failed run starts Auto-Healing automatically. By default the fixes arrive as a PR, and progress is reported as a comment on the originating PR, with no extra wiring on pull_request events. To start heal sessions without auto-creating a PR, add auto-create-pr: false:

Make the workflow wait and pass or fail

By default the action exits as soon as Checksum accepts the run (about 15 seconds), and you hear about results through the PR comment. Set wait: true if the workflow check itself should pass or fail on the outcome. The step succeeds only when the run’s verdict is pass. A failed run, an empty selection, a cancelled run, an infrastructure error, or a timeout all fail the step.
Runner minuteswait: true holds a runner for the whole run, typically 5–25 minutes. If runner cost matters, keep wait: false and rely on the PR-comment notification.

Split a large run across shards

Set shard-count (2–40, in grep and affected modes) to run the tests in parallel and merge the results into one verdict:
Before your first sharded runUpdate checksumai on your tests branch to 4.4.0 or later and commit it (npm install checksumai@latest is the safe default). Older versions can fail to merge shard reports, which delays the final verdict well past a normal run. With wait: true and no wait-timeout-seconds, the step then runs until the job’s own timeout-minutes. Always set wait-timeout-seconds (or a job-level timeout-minutes) alongside wait: true.
Sharding works together with auto-heal from action v2.1.0 onward: once the shards merge, a merged run that ends failed is healed just like a non-sharded run. If the shards never merge, the heal is never evaluated. Older action versions reject the combination; @v2 already resolves to the latest 2.x. Check your suite against Sharding before using a large shard count.

Test each PR’s preview deployment

If every PR is deployed to its own preview URL, pass env-overrides (grep mode only) to point the run at it:

Choose the tests-repo branch

In grep mode, branch picks the branch of your tests repository that the run checks out. Leave it out to use the tests repo’s default branch. Set it explicitly when the workflow runs from your application repo, so the run (and any heal PR, which defaults to the run’s branch) uses your tests repo’s integration branch rather than the application PR’s branch:

Pin the action version

Use @v2 for compatible updates (recommended), @v2.0.0 for a release-specific tag, or a full commit SHA if you need a guarantee of immutability, since tags can technically be moved by the repo owner (see GitHub’s guidance on pinning third-party actions). @v1 still exists with the pre-sharding behavior but no longer gets updates.
Wraps the public-API execution endpoints in a single step with built-in auto-heal on failure. Tests execute in Checksum’s cloud, so no Node, Playwright, or browsers are needed on the runner. Source PR auto-detected: from the event payload on pull_request, via the GitHub API on push.

Required secrets

Action inputs

Set exactly one execution mode: grep, affected, suite-ids, test-ids, or collection-id.

Action outputs

Step exit with wait: true

Version pins

Rules: @v2 fails the step on an empty selection. Sharding requires checksumai ≥ 4.4.0 on the tests branch; older versions can fail to merge shard reports and delay the terminal verdict. Older action versions reject shard-count + auto-heal client-side. From application-code CI, set branch to the tests repo’s integration branch. PR comments come from Checksum’s GitHub App, so workflows need no comment steps and only pull-requests: read.

Run the CLI in GitHub Actions

Use this when the tests should run on your own runner, for example to reach an app that’s only available on your network. Create .github/workflows/checksum-tests.yml in your tests repository:
.github/workflows/checksum-tests.yml
Add the secrets the workflow uses under Settings → Secrets and variables → Actions: CHECKSUM_API_KEY (required), plus the test user’s USERNAME and PASSWORD, your application’s BASE_URL, and its LOGIN_URL. Setting CI: true turns on report uploads and auto-heal PRs by default (see checksum.config.ts). Variables you set explicitly always override the downloaded .env (Environment variables). To run only the tests a PR affects, add --cksm-affected to the test command (Selecting tests).

If your tests live in another repository

When the workflow runs in your application repo but the tests live elsewhere, check out the tests repo with a Personal Access Token stored as TEST_REPO_PAT:

Try the workflow

  1. Commit the workflow file to your main branch.
  2. Go to Actions, select the workflow, and click Run workflow.
  3. Watch the run to confirm tests execute and the report appears in Test Results.
Steps in order: actions/checkout@v4 → npm install → npx playwright install --with-deps → npx checksumai dotenv --download --api-key="${{ secrets.CHECKSUM_API_KEY }}" → npx checksumai test with env vars below. Optional flags: --cksm-affected (PR-affected tests only), --cksm-auto-heal (heal failures).

Secrets for the CLI workflow

CI: true sets defaults hostReports: true and autoHealPRs: true. Explicit env vars take precedence over the downloaded .env.

Run the CLI in GitLab CI/CD

Add this job to your .gitlab-ci.yml:
.gitlab-ci.yml
Add the same variables under your GitLab project’s Settings → CI/CD → Variables. In merge-request pipelines, --cksm-affected picks up the target branch from CI_MERGE_REQUEST_TARGET_BRANCH_NAME automatically.

If your tests live in another project

Clone the tests project with an access token stored as TEST_REPO_PAT:
Image node:20-bookworm. before_script: npm ci → npx playwright install --with-deps → npx checksumai dotenv --download --api-key="${CHECKSUM_API_KEY}". script: npx checksumai test. Set CI: "true". In merge-request pipelines, --cksm-affected reads CI_MERGE_REQUEST_TARGET_BRANCH_NAME.

GitLab CI/CD variables

Location: GitLab project → Settings → CI/CD → Variables.

Trigger and gate a run from any CI system

Any CI system that can run curl can use Checksum. Start a run, check its status until it finishes, and pass the build only when the verdict is "pass". On GitHub, the GitHub Action does exactly this for you and is usually simpler.
  1. Start the run. For PR checks, use POST https://api.checksum.ai/public-api/v2/execution/grep, which accepts the PR branch, a preview URL in envOverrides, and a shardCount in one request. Save the runId it returns. The other ways to start a run are in Execution endpoints.
  2. Check the status with GET https://api.checksum.ai/public-api/v1/execution/status/run/{runId} until isTerminal is true (Run status).
  3. Gate on verdict, not on the passed/failed counts. The verdict is only computed after sharded results merge, and it treats an empty selection as a failure.
  4. Heal, if you like, by including an autoHeal block when you start the run, or by calling POST https://api.checksum.ai/public-api/v1/auto-heal afterward (Auto-Healing).
Here is a complete, sharded PR check. It’s written in GitHub Actions syntax; adapt the ${{ … }} expressions to your CI system’s variables. It needs a CHECKSUM_API_KEY secret, jq on the runner (preinstalled on ubuntu-latest), and a per-PR preview URL.

Example: a sharded PR check with curl

.github/workflows/checksum-sharded.yml
API runs can’t be cancelledCancelling your CI job (for example with concurrency: cancel-in-progress) only stops your polling. The Checksum run continues until it finishes. Plan your concurrency settings with this in mind.
Starts a full-suite run with auto-heal on failure, polls until it finishes, and exits non-zero unless the verdict is pass. When the run fails, Checksum starts healing automatically, so no separate heal call is needed. Needs curl, jq, and CHECKSUM_API_KEY.
To heal a run that already finished without autoHeal, call POST /auto-heal with testRunId set to the run’s UUID (the runId above), not the dispatch name (see the next example).
The per-test results are covered in Results API, and healing in Auto-Healing.
All calls send Authorization: Bearer $CHECKSUM_API_KEY and, with a body, Content-Type: application/json.
  • envOverrides: application variables only. CI and keys starting with CHECKSUM_ are rejected with 400.
  • v1 execution endpoints can shard, but they run against the project’s configured branch and environment. For a PR branch with a preview URL, use v2 grep.
  • POST /v1/execution/suite with {"shardCount": 4} returns { "runId": "…", "name": null, "sharded": true }.
  • POST /v1/auto-heal returns { "batchId", "sessionIds", "failureCount", "testIds" }. Poll GET /v1/auto-heal/batch/{batchId}.
  • API-triggered runs can’t be cancelled through the public API. Cancelling the CI job only stops polling.

Auto-heal tests that fail in CI

Checksum can fix the tests that failed in a CI run and open a PR with the fixes. You can opt in for a single run, or ask Checksum to turn it on for the whole project. For one run, add auto-heal: true to the GitHub Action, add --cksm-auto-heal to the CLI test command, or include an autoHeal block when you start a run through the API. With the CLI in GitHub Actions or GitLab, the repository, branch, and PR number are detected from the CI environment:
The full --cksm-auto-heal* flag set is in Running Tests → test flags.

Heal every failing CI run (project-wide)

Enabled by ChecksumChecksum can turn on automatic healing for your whole project, so every failing CI run is healed without per-pipeline flags. Ask your Checksum team to enable it.
Once it’s enabled, set this variable in each pipeline whose runs should be treated as CI runs:
A per-run opt-in (CLI flag, action input, or API autoHeal) takes priority over the project-wide setting for that run. See Auto-Healing for how healing works.

Per-run opt-in

Project-wide

Enabled by Checksum on request. Then set CHECKSUM_RUNTIME_REQUEST_REASON: cicd in each pipeline whose runs are CI runs. Precedence: a per-run opt-in overrides the project-wide setting for that run.
i
Planning your pipeline
When to run, and how to fix common CI problems

When to run your tests

Start with a manual trigger to verify the setup, then add scheduled or PR/merge-triggered runs. For schedules Checksum runs for you, without your own CI, see Running Tests → Scheduled runs.

Troubleshooting

As of @v2, a grep (or affected set) that matches no tests fails the step. Check your pattern with npx checksumai test --cksm-affected-dry-run or run the grep locally.
The tests branch has a checksumai version older than 4.4.0, so shard reports can’t merge. Upgrade, commit, and re-run. Always bound wait: true with wait-timeout-seconds.
When the pipeline runs from application code, point the run at your tests repo’s integration branch (often main). With the GitHub Action, set the branch input: healing defaults to the run’s branch. With the REST API, set autoHeal.branch; with the CLI, --cksm-auto-heal-branch. See Where healing runs.
Expected behavior. API-triggered runs can’t currently be cancelled through the public API.
Report upload defaults to on only when CI=true. Set CI: true in the job, or set options.hostReports: true in checksum.config.ts.

Running Tests

Execution and status endpoints, test selection, and verdicts.

Sharding

Parallelize large suites safely.

Auto-Healing

What happens after a failing run.

The Checksum CLI

Every checksumai command and flag.

CI Setup Prompts

Have a coding agent build these workflows.