Skip to main content

At a glance

  • Push first: MCP runs in Checksum’s cloud and only sees pushed commits.
  • PRs: a PR-based generate run (pull request number, repository, and branch) and a heal run open a PR by default. A run from a described flow, a detect session, or a targeted generate leaves its changes in the session: pass autoCreatePR: true up front, or call checksum_session_create_pr afterwards.
  • Deep mode: detect, generate, and heal run in Standard mode unless you pass deepMode: true. A deep-mode session pauses for plan approval (checksum_session_approve).
  • Several projects: tools that act on one project need an applicationId. Tell your agent which project you mean.
Developer guide
Connect and use the cloud MCP server

MCP server: connect and use

Checksum runs a hosted MCP server, so any MCP-capable coding agent can drive Checksum directly. Connect once, then ask in plain English:
  • “Generate Checksum tests for this PR.”
  • “My checkout test is failing. Heal it.”
  • “Why did the last test run fail?”
  • “Detect the flows in our billing area and generate tests for them.”
The work runs in Checksum’s cloud and comes back as a pull request. There’s no local test engine and no Playwright install.
New to MCP?MCP is a standard way for AI coding assistants to use outside tools. Once Checksum is connected over MCP, your assistant (Claude Code, Cursor, Copilot) can generate and fix tests for you when you ask. You don’t need to understand the protocol. Connect once, then talk normally.

Before you start

You need all three:
  1. A Checksum project that’s set up, with environment URL, test users, and Git connected. Checksum configures this during onboarding. Without a project, the MCP server has nothing to act on.
  2. An active login at app.checksum.ai in your browser. The approval screen needs it.
  3. An MCP-capable client, such as Claude Code, Cursor, VS Code (Copilot), Claude desktop/web, or any other MCP client.

Connect your client (browser sign-in)

You sign in through your browser and approve which projects the client may use. There’s no API key to copy, paste, or accidentally commit.
Install the Checksum plugin. It adds the connection plus skills that push your branch and find the right pull request before a cloud run:
Claude Code
Then run /mcp, pick checksum, and approve in the browser.
Then run /mcp and pick checksum to sign in.

What you’re approving

The browser shows which projects the client may act on. If you can access several, you can grant only the ones you want. The client can never reach a project you didn’t approve. Access is re-checked on every request, so if you lose access to a project, the client does too. Connections are reviewed and revoked in the web app (Review or revoke MCP connections).

Check that it worked

Ask your agent, in plain English:
Prompt
It replies with the project you’re connected to and a link to its dashboard. That one call tests the whole chain: client, connection, sign-in, and project access. If it doesn’t work, see Troubleshooting.
Push your work firstChecksum generates in the cloud by checking out your repo from your Git provider, so it only sees commits you’ve pushed. If you ask it to cover code that’s still on your machine, it will test the old version. Commit and push, then ask. The Claude Code plugin does this for you.

What you can ask

Once you’re connected, talk to your agent normally. It picks the right tool.
When a PR is openedChecksum opens a pull request when it can. A PR-based generate run (one where it knows the pull request number, the repository, and the branch) and a heal run open one by default. A run from a flow you describe, a detect session, or a targeted generate leaves its changes in the session for review. Ask for a PR up front (autoCreatePR), or once you’ve reviewed it (checksum_session_create_pr).
  • Prerequisites: a Checksum project (set up during onboarding), an active login at https://app.checksum.ai in the same browser, and an MCP-capable client.
  • Approval scope: per project. Access is re-checked on every request. Connections belong to the user and are revoked from My Profile → MCP connections.
  • Verification: the prompt Run checksum_whoami. returns the connected project and a dashboard link.
  • A PR opens by default for PR-based generate runs (PR number + repository + branch known) and heal runs. Other runs open one with autoCreatePR: true or checksum_session_create_pr.

What your agent can do

You don’t need to name tools yourself. Ask for what you want and your agent picks the right one. Behind the scenes, the Checksum MCP server gives your agent these abilities:
  • Check the connection and see which projects you can use.
  • Detect flows worth testing in your app (and optionally its code), written up as test specs.
  • Generate tests for a pull request, for a flow you describe, or for specs you already have.
  • Run tests: the whole suite, a collection, specific tests, or tests matching a pattern.
  • Heal a failed run, opening a pull request with the fixes.
  • Follow a run as it happens, including what the agent is doing, which files it changed, and the PR it opened, and steer it with follow-up instructions.
  • Approve a plan or answer a question a session is waiting on, so a session started from your editor never has to wait on the web app.
  • Stop, restart, or open a PR for a session.
  • Find recent test runs and sessions, and pull results, the HTML report, and a failing test’s trace, screenshots, and video.
Anything that starts a cloud run or opens a pull request is marked as such, so a good client asks you to confirm first. Every result includes links into the Checksum web app, so you can open a session or run directly instead of working with IDs. Healing doesn’t hide real problems: if a test fails because your app is broken, it reports a real bug instead of rewriting the test to pass. Report and trace links are signed and expire, so they don’t stay live forever in a chat log.

Deep mode and the approval loop

Detect, generate, and heal run in Standard mode by default: fully autonomous, start to finish. Pass deepMode: true for Deep mode: slower, more thorough, and it pauses after planning for your approval. Over MCP that loop is explicit:
  1. Poll checksum_session_status until nextAction is "approve".
  2. Read planMd, the plan awaiting approval. planMdTruncated tells you if it was cut short; the full plan is in the web app.
  3. Call checksum_session_approve to continue into implementation. To ask for changes instead, send them with checksum_session_prompt. A prompt during approval reopens the planning step; it does not approve it. Poll again for the revised plan.
Questions work the same way. When nextAction is "answer", read pendingQuestions and call checksum_session_answer with answers: one array per sub-question of the first pending group, in order, each holding the selected option labels (exactly one unless the sub-question allows multiple; free text is accepted when no option fits). For example, for two sub-questions: [["Playwright"], ["Login", "Checkout"]]. Pass questionId to address a specific group when more than one is pending.
Approval is a deliberate gateThere’s no auto-approve flag. If nobody will be around to review the plan, leave deep mode off.

Typical flows

Detect flows, then generate tests for them. Ask: “Detect the user flows in the billing area, then generate tests for them and open a PR.” Your agent calls checksum_detect with collectionNames: ["Billing"], polls checksum_session_status by sessionId until nextAction is "done", then calls checksum_test_generate with the same collectionNames (or the userStoryIds it detected) and autoCreatePR: true. Start a run and read the results. Ask: “Run the Checkout collection against the preview deployment and tell me what failed.” Your agent calls checksum_test_run (mode: "collection"), waits for the run to finish via checksum_test_run_list, then calls checksum_test_run_download, with a testId for the trace and screenshots of a failing test. Heal a failing run and open the PR. Ask: “Heal the failures in the last Checksum run.” Your agent calls checksum_test_heal with the testRunId, polls checksum_session_status by batchId until it’s done, and hands you the prUrl. If a session finished without a PR, it calls checksum_session_create_pr.

Working with several projects (applicationId)

If your account can access more than one Checksum project, the tools that act on a single project need to know which one. Otherwise you’ll see:
Tell your agent which project you mean, and it will pass the right ID. Run checksum_whoami to see the list. An API key is always tied to exactly one project, so this doesn’t happen with API keys.
Tools marked “Changes things” start billable cloud runs and sessions, steer them, and can open pull requests. Clients should confirm before calling them.

checksum_whoami: check the connection

One call tests the whole chain: client, connection, sign-in, and project access. Run it first when anything looks wrong, and to see the project list when you work with several projects.

checksum_detect: discover flows to test

Type: changes things. Returns: a sessionId. Poll checksum_session_status with it. When the session finishes, the specs are ready for checksum_test_generate (by collection or by userStoryIds), and autoCreatePR or checksum_session_create_pr puts them in a pull request. See also Detect Test Flows.

checksum_test_generate: start test generation

Give it one of three inputs: a pull request, a description, or specs you already have.Type: changes things. Returns: a batchId (a sessionId for a targeted run). Asking twice for the same pull request won’t start a second run. The request joins the run already in progress. Push your commits first, because cloud generation only sees pushed code. See also Generate Tests.

checksum_test_heal: heal a failed run

Type: changes things. Returns: a batchId. It opens one heal session covering all the failing tests in that run and, by default, a pull request with the fixes. Healing works from a run that has already finished, so there’s nothing to push first.Healing doesn’t hide real problems. If a test fails because your app is broken, it reports a real bug instead of rewriting the test to pass. Healing a run with nothing to heal returns the reason:
Pass deepMode: true for Deep mode. See also Auto-Healing.

checksum_test_run: start a test run

Type: changes things. Returns: the testRunId. Read the results with checksum_test_run_download once checksum_test_run_list shows the run finished. See also Running Tests.

checksum_session_status: watch a run

Type: read-only. Your agent polls it until the run finishes. It reports overall progress, the agent’s latest messages, the files it changed, and the prUrl it opened. That lets your agent tell you what’s happening instead of just “still running.”Poll by batchId for a generate or heal run, or by sessionId alone for a detect or targeted-generate session.

checksum_session_prompt: steer a session

Type: changes things. Sends a follow-up instruction into a session, the same as typing into the session in the web app. It works on running, waiting, paused, and completed sessions, and resumes a stopped one. A session that was cancelled, or failed without a way back, won’t accept a prompt, and you’ll get an error saying so. During plan approval, a prompt reopens planning instead of approving.

checksum_session_approve: approve a plan

Type: changes things. Approves the plan a deep-mode session is waiting on (nextAction is approve), so it continues into implementation. See Deep mode and the approval loop.

checksum_session_answer: answer the agent

Type: changes things. Use it when nextAction is answer.

checksum_session_stop and checksum_session_restart

Type: changes things. checksum_session_stop pauses a session and keeps its workspace. Resume it with checksum_session_prompt, or start over with checksum_session_restart. A restart is a new session (the original is kept as a record), so poll the returned restartedAiAgentsSessionId from then on. Restart works on stopped, failed, and finished sessions.

checksum_session_create_pr: open the pull request

Type: changes things. Opens the pull request for a completed session that hasn’t opened one yet, and returns the existing prUrl if it already has one. The call is synchronous and can take a minute or two. If it times out, poll checksum_session_status for the prUrl or call it again.

checksum_session_list: list sessions

Type: read-only. Lists recent agent sessions.

checksum_test_run_list: list test runs

Type: read-only. Lists recent test runs, newest first, so the agent can find a failing run (for example, a testRunId for checksum_test_heal) without you hunting for an ID.

checksum_test_run_download: pull results, reports, and traces

Type: read-only. Report and trace links are signed and expire, so they don’t stay live forever in a chat log. It lists your standard end-to-end runs. API-test runs and hidden runs don’t appear.

MCP: connect with an API key

Browser sign-in needs a browser. For CI, a container, or a client that only supports static tokens, use an API key instead. If browser sign-in worked, skip this. Get the key from Settings → Project Settings in the web app. It’s the same key the CLI and CI use (see API Keys). Send it as a bearer token:
Key stored in plain textYour shell expands the value when you run this, so the key is written into Claude Code’s config file as plain text. Treat that file as a secret, or use browser sign-in, which stores no key at all.
An API key can start real cloud runs and open pull requests on your repo. Keep it out of any file you might commit, and use browser sign-in whenever you have a browser.

MCP troubleshooting

Most clients only read MCP config at startup, so restart the client after editing mcp.json.Then ask your agent to run checksum_whoami. If that works, the connection is fine and the problem is your client’s tool list, not Checksum.
If you signed in through the browser, the session may have expired. Reconnect from your client’s MCP settings.If you’re using an API key, check that it’s sent as Authorization: Bearer <key> and that it hasn’t been rotated in Settings → Project Settings.If you set up an API key earlier and are now switching to browser sign-in, remove the old entry and its Authorization header first. A stale header takes precedence over the browser session, and the client won’t fall back.
Approval needs an active Checksum web-app session. Log in to app.checksum.ai in the same browser, then start the connection again.If you have no projects yet, or yours is still pending review, there’s nothing to authorize. Your Checksum team sets up the project during onboarding.
On a corporate network or VPN, traffic to api.checksum.ai may be blocked or intercepted. Check that you can reach it:
If that hangs or fails, ask IT to allow api.checksum.ai over HTTPS (port 443). A TLS-inspecting proxy can also break the connection by re-signing certificates. It needs api.checksum.ai on its bypass list. See Network access.A proxy configured in your shell isn’t visible to a GUI app like Cursor or VS Code unless you launch the app from that shell.
Checksum checks out your repository from the remote, so it only sees pushed commits. If your work is still local, commit and push, then ask again.
Checksum opens a PR by default only when it has the pull request number, the repository, and the branch. A run from a plain description, a detect session, or a targeted generate leaves its changes in the session. Ask your agent to open one with checksum_session_create_pr, or pass autoCreatePR: true next time.
Healing only works on runs that actually failed, and the message says why this one was skipped. Ask your agent to list recent runs and pick one with failures.
Ask your agent to check checksum_session_status. If nextAction is approve, the session is a deep-mode run waiting on its plan: approve it with checksum_session_approve. If it’s answer, answer the pending question with checksum_session_answer. See Deep mode and the approval loop.
Checksum uses Streamable HTTP, not SSE. Use the same URL and choose the HTTP transport.A client that can only launch local (stdio) servers can’t connect directly. Most such clients can reach a remote server through a bridge, but the simplest option is to use a supported client: Claude Code, Cursor, VS Code (Copilot), or Claude web and desktop (Connect your client).
▦
In the Checksum web app
Review and revoke MCP connections

Review or revoke MCP connections

Review or revoke a connection any time from the MCP connections card on My Profile in the web app. A connection belongs to you, not to a project, so revoking it cuts the client off from every project you granted. If you connect with an API key instead, the key lives in Settings → Project Settings (see API keys).
i
How it works
MCP security and related pages

MCP security

  • Your code is never uploaded from your machine. Checksum clones your repository from your connected Git provider, in the cloud. The MCP server only carries instructions and results.
  • Browser sign-in is scoped. You choose which projects a client may touch, and access is re-checked on every request. Revoke it any time from My Profile.
  • Report and trace links are signed and expire, so a URL pasted into a chat log doesn’t stay live forever.
  • Tools that cost money, steer sessions, or open PRs are marked as such, so your client can ask before running them.

Generate Tests

Every generation trigger, including MCP.

Auto-Healing

What healing does to a failing run.

API Keys & Authentication

Where the key comes from and how to keep it safe.

CI Setup Prompts

Prompts that have your agent build Checksum GitHub Actions workflows.