> ## Documentation Index
> Fetch the complete documentation index at: https://checksum.ai/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Sharding

> Sharding splits one run across up to 40 parallel shards and merges their results into a single report and verdict. It's the fastest way to cut wall-clock time on a large suite, but only if your tests are built to run in parallel. This page covers how to turn it on, how to tell whether your suite is ready, and the current limits.

## At a glance

| Step | How | Watch out for |
| - | - | - |
| [Update the runtime](#update-checksumai-first) | `checksumai` 4.4.0 or later on the tests branch | Without it, shard results never merge |
| [GitHub Action](#shard-a-run-from-the-github-action) | Add `shard-count` (2–40) | Pair `wait: true` with a timeout |
| [REST API](#shard-a-run-from-the-rest-api) | Add `shardCount` (2–40) to the run request | Track the run by `runId` |

<div className="ai-ref">
  <Accordion title="Reference for AI: sharding at a glance" icon="robot">
    | Interface | Setting | Allowed values | Key constraint |
    | - | - | - | - |
    | CLI / tests branch | `npm install checksumai@latest`, committed to the tests branch | `4.4.0` or later | Required before the first sharded run, or results never merge |
    | GitHub Action | `shard-count` input | `2`–`40` | Honored in `grep` and `affected` modes. With `auto-heal`, requires action v2.1.0+. |
    | REST API | `shardCount` body field on any execution request | Omit or `1` = not sharded; `2`–`40` = sharded | Values above 40 are rejected, not capped |
    | REST API | `GET https://api.checksum.ai/public-api/v1/execution/status/run/{runId}` | — | Track sharded runs by `runId`. `name` is `null` when sharded. |

    * One Playwright worker per shard: total parallelism is controlled by the shard count alone.
    * `verdict` is computed only after the shard reports merge. Gate CI on `verdict`, not on counts.
    * Auto-heal works with sharding: a merged run that ends `failed` is healed like a non-sharded run.
    * API-triggered runs, sharded or not, can't be cancelled through the public API.
    * Tests must own their data and sessions (see Is your suite ready to shard?, `#suite-readiness`).
  </Accordion>
</div>

<div className="part dev"><span className="part-icon">{"</>"}</span><div><div className="part-title">Developer guide</div><div className="part-sub">Upgrade the runtime, then enable sharding from the GitHub Action or the REST API</div></div></div>

## Update `checksumai` first

Use sharding when a suite is too slow as a single run and you want faster CI feedback. Before your first sharded run, update the Checksum runtime on your tests branch. After all shards finish, Checksum merges their results using the `checksumai` version installed **on the branch you run against**, and the minimum version supported in production for sharding is **4.4.0**.

```bash theme={null}
npm install checksumai@latest
git commit -am "Update checksumai for sharded runs" && git push
```

<Warning>
  **Required before your first sharded run**

  With an older version, the shards still run but their results are never combined. The run never returns a final `verdict`, and any `autoHeal` request is never evaluated. If you're not sure which version you need, use the latest stable release or contact your Checksum team.
</Warning>

## Shard a run from the GitHub Action

Add `shard-count` to `checksum-ai/test-run-action@v2`. It works in `grep` and `affected` modes, with 2 to 40 shards:

```yaml theme={null}
- uses: checksum-ai/test-run-action@v2
  with:
    api-key: ${{ secrets.CHECKSUM_API_KEY }}
    grep: '@smoke'
    shard-count: 8
    wait: true
    wait-timeout-seconds: 1800
```

If you use `wait: true`, always pair it with `wait-timeout-seconds` (or a job `timeout-minutes`). Otherwise an outdated `checksumai` on the tests branch would keep the step waiting until the job times out. To combine sharding with `auto-heal`, you need action v2.1.0 or later; `@v2` already resolves to the latest 2.x. All the action's inputs are in [CI/CD Integration → GitHub Action](/docs/ci-integration#run-checksum-with-the-github-action).

<div className="ai-ref">
  <Accordion title="Reference for AI: shard-count in the GitHub Action" icon="robot">
    | Input | Value | Why |
    | - | - | - |
    | `shard-count` | `2`–`40` | Number of parallel shards. Honored in `grep` and `affected` modes only. |
    | `wait` + `wait-timeout-seconds` | e.g. `true` + `1800` | Always pair `wait: true` with `wait-timeout-seconds` (or a job `timeout-minutes`). With an outdated `checksumai`, the step would otherwise wait until the job timeout. |
    | `auto-heal` | `true` | `shard-count` + `auto-heal` requires action **v2.1.0+**. Older versions reject the combination client-side. `@v2` resolves to the latest 2.x. |
  </Accordion>
</div>

## Shard a run from the REST API

Add `shardCount` to the body of any request that starts a run: `POST /public-api/v1/execution/suite`, `/execution/collection/{id}`, `/execution/tests`, or `POST /public-api/v2/execution/grep` (see [Execution endpoints](/docs/running-tests#start-a-cloud-run-from-the-rest-api)). Leave it out, or set `1`, for a normal run. Set 2 to 40 to run in parallel.

### Shard the full suite

```bash theme={null}
curl -X POST https://api.checksum.ai/public-api/v1/execution/suite \
  -H "Authorization: Bearer $CHECKSUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"shardCount": 4}'
# → { "runId": "9f2c7a4e-8b31-4d6a-a2f0-3c5e1b7d9a42", "name": null, "sharded": true }
```

A sharded run has no job `name` (it comes back as `null`), so always track it by `runId`.

### Shard a PR branch against its preview URL

For pull-request checks, use v2 grep, which lets you pick the branch and override the URL in the same request. The v1 endpoints shard too, but they always run against the project's configured branch and environment.

```bash theme={null}
curl -X POST https://api.checksum.ai/public-api/v2/execution/grep \
  -H "Authorization: Bearer $CHECKSUM_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "grep": "@smoke",
    "branch": "feature/checkout-preview",
    "envOverrides": { "BASE_URL": "https://pr-42.preview.example.com" },
    "shardCount": 8
  }'
```

A full PR-gating workflow is in [CI/CD Integration → sharded PR check](/docs/ci-integration#example-a-sharded-pr-check-with-curl).

### Follow a sharded run

Check the run by its ID until `isTerminal` is `true`, then pass or fail on `verdict`. While it runs, `phase` moves from `sharding` to `merging` to `complete`.

```bash theme={null}
curl https://api.checksum.ai/public-api/v1/execution/status/run/$RUN_ID \
  -H "Authorization: Bearer $CHECKSUM_API_KEY"
```

```json theme={null}
{
  "status": "running", "isTerminal": false, "verdict": "pending",
  "phase": "sharding", "executedCount": 0,
  "sharded": true, "shardTotal": 8,
  "passed": 0, "failed": 0, "recovered": 0, "bug": 0, "failureReason": null
}
```

The full field list is in [Running Tests → Run status](/docs/running-tests#check-whether-a-run-passed). The older job-name status endpoint doesn't support sharded runs.

<div className="ai-ref">
  <Accordion title="Reference for AI: shardCount in the REST API" icon="robot">
    Accepted on `POST https://api.checksum.ai/public-api/v1/execution/suite`, `POST https://api.checksum.ai/public-api/v1/execution/collection/{id}`, `POST https://api.checksum.ai/public-api/v1/execution/tests`, and `POST https://api.checksum.ai/public-api/v2/execution/grep`. Headers: `Authorization: Bearer $CHECKSUM_API_KEY`, `Content-Type: application/json`.

    | Field | Type | Required | Description |
    | - | - | - | - |
    | `shardCount` | integer | No | Number of parallel shards. Omit it, or set `1`, for a non-sharded run. Set `2`–`40` to run in parallel. Values above 40 are rejected with `400` (not capped). Each shard runs one Playwright worker. |

    | Response field | Sharded value |
    | - | - |
    | `runId` | UUID. Use for status polling. |
    | `name` | `null` |
    | `sharded` | `true` |

    Status: `GET https://api.checksum.ai/public-api/v1/execution/status/run/{runId}`. `phase`: `sharding` → `merging` → `complete`. Poll until `isTerminal` is `true`; gate on `verdict` (computed only after the merge). Sharded status fields: `sharded: true`, `shardTotal`. The legacy job-name status endpoint doesn't support sharded runs. Use v2 grep for PR-scoped sharded runs (`branch`, `envOverrides`); v1 endpoints run against the project's configured branch and environment. Prerequisite: `checksumai` ≥ 4.4.0 on the branch being run.
  </Accordion>
</div>

<div className="part bg"><span className="part-icon">i</span><div><div className="part-title">How it works</div><div className="part-sub">The shard lifecycle, suite readiness, limits, and rollout</div></div></div>

## How sharding works

<div className="flow">
  <div className="node"><b>Dispatch</b><span>One request with <code>shardCount: N</code></span></div>
  <div className="arrow">→</div>
  <div className="node"><b>Shard</b><span>N machines, one Playwright worker each, run in parallel</span></div>
  <div className="arrow">→</div>
  <div className="node"><b>Merge</b><span>Shard reports combine into one report</span></div>
  <div className="arrow">→</div>
  <div className="node"><b>Verdict</b><span>A single <code>pass</code> / <code>fail</code>, plus auto-heal if you opted in</span></div>
</div>

* **Parallelism is controlled only by the shard count.** Each shard always runs one Playwright worker.
* **Status reporting is unified.** Poll `GET https://api.checksum.ai/public-api/v1/execution/status/run/{runId}`. Its `phase` moves through `sharding` → `merging` → `complete`, and `verdict` is computed only after the merge (see [Run status](/docs/running-tests#check-whether-a-run-passed)). The legacy job-name status endpoint doesn't support sharded runs.
* **Auto-heal works with sharding.** Once the shards merge, a merged run that ends `failed` is healed exactly like a non-sharded run. A lost shard doesn't stop the surviving shards' failures from being healed.

## Is your suite ready to shard?

A suite written to run serially, or with a few local workers, can hit new failure modes once it's split across many machines hitting your application at the same time. Check these first.

### Every test owns its data

Data isolation is the single biggest predictor of a clean sharded run. Shards run independently and in parallel against the same environment, so:

* When the setup isn't what the test is verifying, create the data through your API rather than a UI form, with a **unique name** like `<test-id>-<random-suffix>`. Never use a fixed human-readable name another test might also create or match.
* Don't assert on the state of a shared resource ("there are 3 items in the list"). Assert on your own uniquely named item.
* Avoid a single default project, workspace, or org that every test reads and writes. If several tests mutate the same record, one shard's setup or cleanup can race another shard's assertions.
* Clean up what you created in a best-effort teardown step. Log a cleanup failure rather than failing the test, so leftover data doesn't go unnoticed.

<Tip>
  **Moving setup from UI to API?**

  Confirm the API call creates *everything* the test depends on, not just a top-level record. If a setup call only creates a parent object while the test also depends on nested state, you won't save the time you expect, and part of the real setup may still go through the slow UI path.
</Tip>

### Session and login state

If your app keeps session- or user-scoped state that a test's steps depend on, each shard logging in through the UI may need its own isolated user or session. Otherwise two shards logged in as the same user can race each other's navigation, toggles, or in-progress work.

* **Measure before you assume.** Check how many concurrent sessions your backend actually supports on one account. Some backends safely deduplicate concurrent logins into one valid session, while others invalidate the previous session on each new login. Open N sessions concurrently and confirm none get logged out before you pick a shard count. Add test users per role in [Environments](/docs/environments#test-users).
* **Restore saved sessions early.** If you restore cookies or local storage instead of logging in through the UI, apply them **before** the first navigation, and don't rewrite cookie domains you don't need to. Some SSO providers run a silent login check on page load that can overwrite a session restored too late or with a mangled cookie domain. This shows up as an intermittent "logged out" failure that only happens in parallel runs.

### Backend and environment capacity

Every shard's tests hit your application's APIs at the same time:

* Check for rate limits, connection-pool limits, or per-account concurrency limits on your backend and staging/test environment.
* If you only see timeouts in sharded runs, **don't raise the timeout first**. Rule out a data or session race (two shards fighting over the same entity or account), because it fails the same way real saturation does. If the failure still reproduces with isolated data, it's a capacity signal, and raising the timeout is the right fix.
* Retries add wall-clock time to every retried test. On a suite with a few very slow tests, they can cancel out a meaningful part of the speedup.

### Avoid `test.describe.serial` as a workaround

If two tests only pass when forced to run in order, they usually share state that should be isolated. Forcing serial execution hides the coupling and its runtime cost instead of fixing it. It doesn't help sharding either, because tests locked into a fixed order gain nothing from more shards.

## Know your ceiling

* **A shard's runtime depends on the files assigned to it.** If one file is much larger or slower than the rest, its shard sets your minimum wall-clock time no matter how many shards you add. Splitting an oversized file into smaller files gives sharding more to balance.
* **The merge takes time.** Once every shard finishes, Checksum combines the reports. Budget a short merge step on top of shard runtime.
* **Compare runs fairly.** A shard-count change is only comparable to an earlier run if the test count, environment, and retry settings also match.
* **Verify in real CI.** Auth and session handling can behave differently on a headless CI runner than locally. Confirm with an actual sharded run before you rely on a shard count.

## Current limits

| Limit | Value |
| - | - |
| Shard count | `2`–`40` per run |
| Workers per shard | Always `1`. Total parallelism is controlled by shard count alone. |
| Requests above the maximum | Rejected outright, not capped to 40 |
| Minimum `checksumai` | `4.4.0` on the branch being run |
| Auto-heal + sharding | Supported. The merged run is healed like a non-sharded run. Requires action v2.1.0+ when using the GitHub Action. |
| Cancellation | API-triggered runs (sharded or not) can't be cancelled through the public API |

Need more than 40 shards? [Contact your Checksum team](mailto:support@checksum.ai).

## Roll out safely

<Steps>
  <Step title="Upgrade">
    Commit `checksumai` ≥ 4.4.0 to the tests branch.
  </Step>

  <Step title="Start small">
    Run with a modest shard count (e.g. 2–4) and confirm it's fully green and stable.
  </Step>

  <Step title="Diagnose every new failure">
    Classify it as isolation or capacity (see [Is your suite ready to shard?](#is-your-suite-ready-to-shard)) before you reach for a timeout, a retry, or more shards.
  </Step>

  <Step title="Scale gradually">
    Increase the shard count step by step and compare wall-clock times fairly.
  </Step>
</Steps>

## Troubleshooting

<AccordionGroup>
  <Accordion title="The run stays in merging / never returns a verdict">
    The tests branch has `checksumai` older than 4.4.0. Upgrade, commit, and re-run.
  </Accordion>

  <Accordion title="400 error when setting shardCount">
    Values above 40 are rejected, not capped. Use `2`–`40`, or `1`/omit for non-sharded.
  </Accordion>

  <Accordion title="Intermittent &#x22;logged out&#x22; failures only when sharded">
    This is usually a session race or a session restored too late. See [Session and login state](#session-and-login-state).
  </Accordion>

  <Accordion title="name is null in the dispatch response">
    Expected for sharded runs. Poll by `runId` with `GET https://api.checksum.ai/public-api/v1/execution/status/run/{runId}`.
  </Accordion>
</AccordionGroup>

## Related

<CardGroup cols={2}>
  <Card title="Running Tests" icon="play" href="/docs/running-tests">
    Execution endpoints and run status.
  </Card>

  <Card title="CI/CD Integration" icon="code-branch" href="/docs/ci-integration">
    Sharded PR checks end to end.
  </Card>

  <Card title="Story & Test Format" icon="list-check" href="/docs/story-and-test-format">
    Data setup and cleanup in generated tests.
  </Card>
</CardGroup>
