At a glance
Reference for AI: sharding at a glance
Reference for AI: sharding at a glance
- One Playwright worker per shard: total parallelism is controlled by the shard count alone.
verdictis computed only after the shard reports merge. Gate CI onverdict, not on counts.- Auto-heal works with sharding: a merged run that ends
failedis healed like a non-sharded run. - API-triggered runs, sharded or not, can’t be cancelled through the public API.
- Tests must own their data and sessions (see Is your suite ready to shard?,
#suite-readiness).
Update checksumai first
Use sharding when a suite is too slow as a single run and you want faster CI feedback. Before your first sharded run, update the Checksum runtime on your tests branch. After all shards finish, Checksum merges their results using the checksumai version installed on the branch you run against, and the minimum version supported in production for sharding is 4.4.0.
Shard a run from the GitHub Action
Addshard-count to checksum-ai/test-run-action@v2. It works in grep and affected modes, with 2 to 40 shards:
wait: true, always pair it with wait-timeout-seconds (or a job timeout-minutes). Otherwise an outdated checksumai on the tests branch would keep the step waiting until the job times out. To combine sharding with auto-heal, you need action v2.1.0 or later; @v2 already resolves to the latest 2.x. All the action’s inputs are in CI/CD Integration → GitHub Action.
Reference for AI: shard-count in the GitHub Action
Reference for AI: shard-count in the GitHub Action
Shard a run from the REST API
AddshardCount to the body of any request that starts a run: POST /public-api/v1/execution/suite, /execution/collection/{id}, /execution/tests, or POST /public-api/v2/execution/grep (see Execution endpoints). Leave it out, or set 1, for a normal run. Set 2 to 40 to run in parallel.
Shard the full suite
name (it comes back as null), so always track it by runId.
Shard a PR branch against its preview URL
For pull-request checks, use v2 grep, which lets you pick the branch and override the URL in the same request. The v1 endpoints shard too, but they always run against the project’s configured branch and environment.Follow a sharded run
Check the run by its ID untilisTerminal is true, then pass or fail on verdict. While it runs, phase moves from sharding to merging to complete.
Reference for AI: shardCount in the REST API
Reference for AI: shardCount in the REST API
POST https://api.checksum.ai/public-api/v1/execution/suite, POST https://api.checksum.ai/public-api/v1/execution/collection/{id}, POST https://api.checksum.ai/public-api/v1/execution/tests, and POST https://api.checksum.ai/public-api/v2/execution/grep. Headers: Authorization: Bearer $CHECKSUM_API_KEY, Content-Type: application/json.GET https://api.checksum.ai/public-api/v1/execution/status/run/{runId}. phase: sharding → merging → complete. Poll until isTerminal is true; gate on verdict (computed only after the merge). Sharded status fields: sharded: true, shardTotal. The legacy job-name status endpoint doesn’t support sharded runs. Use v2 grep for PR-scoped sharded runs (branch, envOverrides); v1 endpoints run against the project’s configured branch and environment. Prerequisite: checksumai ≥ 4.4.0 on the branch being run.How sharding works
shardCount: Npass / fail, plus auto-heal if you opted in- Parallelism is controlled only by the shard count. Each shard always runs one Playwright worker.
- Status reporting is unified. Poll
GET https://api.checksum.ai/public-api/v1/execution/status/run/{runId}. Itsphasemoves throughsharding→merging→complete, andverdictis computed only after the merge (see Run status). The legacy job-name status endpoint doesn’t support sharded runs. - Auto-heal works with sharding. Once the shards merge, a merged run that ends
failedis healed exactly like a non-sharded run. A lost shard doesn’t stop the surviving shards’ failures from being healed.
Is your suite ready to shard?
A suite written to run serially, or with a few local workers, can hit new failure modes once it’s split across many machines hitting your application at the same time. Check these first.Every test owns its data
Data isolation is the single biggest predictor of a clean sharded run. Shards run independently and in parallel against the same environment, so:- When the setup isn’t what the test is verifying, create the data through your API rather than a UI form, with a unique name like
<test-id>-<random-suffix>. Never use a fixed human-readable name another test might also create or match. - Don’t assert on the state of a shared resource (“there are 3 items in the list”). Assert on your own uniquely named item.
- Avoid a single default project, workspace, or org that every test reads and writes. If several tests mutate the same record, one shard’s setup or cleanup can race another shard’s assertions.
- Clean up what you created in a best-effort teardown step. Log a cleanup failure rather than failing the test, so leftover data doesn’t go unnoticed.
Session and login state
If your app keeps session- or user-scoped state that a test’s steps depend on, each shard logging in through the UI may need its own isolated user or session. Otherwise two shards logged in as the same user can race each other’s navigation, toggles, or in-progress work.- Measure before you assume. Check how many concurrent sessions your backend actually supports on one account. Some backends safely deduplicate concurrent logins into one valid session, while others invalidate the previous session on each new login. Open N sessions concurrently and confirm none get logged out before you pick a shard count. Add test users per role in Environments.
- Restore saved sessions early. If you restore cookies or local storage instead of logging in through the UI, apply them before the first navigation, and don’t rewrite cookie domains you don’t need to. Some SSO providers run a silent login check on page load that can overwrite a session restored too late or with a mangled cookie domain. This shows up as an intermittent “logged out” failure that only happens in parallel runs.
Backend and environment capacity
Every shard’s tests hit your application’s APIs at the same time:- Check for rate limits, connection-pool limits, or per-account concurrency limits on your backend and staging/test environment.
- If you only see timeouts in sharded runs, don’t raise the timeout first. Rule out a data or session race (two shards fighting over the same entity or account), because it fails the same way real saturation does. If the failure still reproduces with isolated data, it’s a capacity signal, and raising the timeout is the right fix.
- Retries add wall-clock time to every retried test. On a suite with a few very slow tests, they can cancel out a meaningful part of the speedup.
Avoid test.describe.serial as a workaround
If two tests only pass when forced to run in order, they usually share state that should be isolated. Forcing serial execution hides the coupling and its runtime cost instead of fixing it. It doesn’t help sharding either, because tests locked into a fixed order gain nothing from more shards.
Know your ceiling
- A shard’s runtime depends on the files assigned to it. If one file is much larger or slower than the rest, its shard sets your minimum wall-clock time no matter how many shards you add. Splitting an oversized file into smaller files gives sharding more to balance.
- The merge takes time. Once every shard finishes, Checksum combines the reports. Budget a short merge step on top of shard runtime.
- Compare runs fairly. A shard-count change is only comparable to an earlier run if the test count, environment, and retry settings also match.
- Verify in real CI. Auth and session handling can behave differently on a headless CI runner than locally. Confirm with an actual sharded run before you rely on a shard count.
Current limits
Roll out safely
Upgrade
checksumai ≥ 4.4.0 to the tests branch.Start small
Diagnose every new failure
Scale gradually
Troubleshooting
The run stays in merging / never returns a verdict
The run stays in merging / never returns a verdict
checksumai older than 4.4.0. Upgrade, commit, and re-run.400 error when setting shardCount
400 error when setting shardCount
2–40, or 1/omit for non-sharded.Intermittent "logged out" failures only when sharded
Intermittent "logged out" failures only when sharded
name is null in the dispatch response
name is null in the dispatch response
runId with GET https://api.checksum.ai/public-api/v1/execution/status/run/{runId}.