Replay testing for AI agents

Test your next prompt change on last week’s real conversations

Tapeback is a testing tool for AI agents. Before a prompt, model or tool change ships, it replays a sample of last week’s real conversations against it, and two models group the steps whose meaning changed, for a person to mark.

Replays run on traffic you already recorded. Nothing is sent to your users.

  • 89%of teams in LangChain’s 2025 agent survey have observability
  • 52.4%run offline evaluations, same survey
  • 7 daysof your own real conversations, sampled for each replay
A group of 212 replayed conversations where the changed agent looks up an order instead of asking for the order number, waiting to be marked.

npx tapeback init

a week later, a pull request edits prompts/refund.md

npx tapeback replay --change prompts/refund.md --days 7

The refund fix, replayed

Tuesday’s refund fix, run against last week before it shipped

I shipped a refund-prompt fix on a Tuesday. It passed every saved case. The agent stopped asking for an order number and answered from the wrong order for two days before anyone saw it. These steps replay that change against the week before, in an example workspace.

A pull request check in CI, in plain words: 1,840 of 12,406 tapes replayed, 1,611 behaved the same, 229 changed in 3 groups, check blocked on 2 unmarked groups.
  1. 01recordOne command wraps your agent, and every conversation becomes a tape.
  2. 02sampleThe refund prompt changes. 1,840 of the week’s 12,406 tapes are drawn.
  3. 03replayThe changed agent gets each recorded history and writes the next step only.
  4. 04group229 steps changed, in three groups. 212 skip the order-number question.
  5. 05markThe pull request stays blocked until a named person marks every group, and writes why.

What sorts the changed steps

Rules compare structure. Two models read what the words now mean.

Their labels sort the report and never clear a change. Only a named person’s mark unblocks the pull request.

01

Embedding model - Reads each changed step and groups the ones that changed alike.

02

Pair classifier - Sets the recorded step beside your changed agent’s step and returns one of six labels, with a confidence.

03

Under 0.80 - No label is shown, and the step waits in Needs a look for an engineer to label by hand.

04

Training - Every mark leaves a labeled pair. Only pairs from opted-in workspaces may train the classifier, never tapes.

Why nothing alerted

A wrong answer and a right one return the same status code.

Nothing alerted, because nothing was broken except the answer.

—Tapebackon what a status code can’t tell you

A decision I’d defend

The check has no pass rate. A person reads the groups and signs them.

A prompt change is supposed to change behavior. So a percentage of changed steps can’t tell a fix from a break.

In the refund run, the fix and the missing question arrive together. Grouping turns 229 changed steps into three groups. A person can read three.

Every agent change waits for a human, including at six on a Friday, and there’s no single number for a dashboard.

  • Tool calls are answered from the tape. A brand-new tool stops the replay there.
  • No simulated user, so a break two turns later stays invisible.
  • Every mark keeps a name and a reason, at a stable URL.

Once the new agent says something different, the real user’s next message no longer answers it.

A model playing your customer is an invented customer.

—Tapebackon why no replay simulates a user

Plugs into your stack

It wraps the framework your agent already runs on

The check runs in GitHub Actions, and a run summary posts to Slack.

npx tapeback import --days 7

  • Vercel AI SDK

    The recorder wraps your agent’s entry point.

  • LangChain

    Same wrapper. Tool arguments and results stay in order.

  • OpenAI Agents SDK

    On OpenAI or Azure OpenAI, init adds the recorder and writes tapeback.config.ts.

  • Claude Agent SDK

    One tape per conversation, tool calls included.

  • OpenTelemetry

    Any trace carrying full message and tool bodies.

  • Langfuse and LangSmith

    Import last week’s traces. Sampled ones are skipped.

Pricing

Team is $240 a month. Raindrop Pro starts at $299, and you’ll still want it.

  • Bench

    $0per month

    One engineer checking one agent’s changes from a laptop.

    • 1 agent, 1 seat
    • 300 replayed conversations a month
    • 7 days of tapes kept
    • CLI replay and the web run report
    • Recording and redaction rules, never metered
    • At the cap, replays pause until the 1st; recording does not
  • TeamCheck on every PR

    $240per month

    A team that ships agent changes through pull requests and wants the check on every one.

    • 3 agents, 10 seats
    • 10,000 replayed conversations a month, then $0.01 each
    • 30 days of tapes kept
    • GitHub check that blocks on unmarked groups
    • Import from Langfuse, LangSmith or OpenTelemetry, never charged
    • Run summary posted to a Slack channel
    • Weekly baseline run and same-commit re-runs, not counted
  • Org

    $900per month

    Several agents, a security review to pass, and a record of who marked what.

    • Unlimited agents, 40 seats
    • 60,000 replayed conversations a month, then $0.008 each
    • 90 days of tapes kept
    • SSO / SAML included
    • Audit log of every mark, rule change and export
    • Signed data processing agreement and sub-processor list
    • Email support answered by the next business day
Get an API key
  • Bench stays free. No end date, no payment details.
  • Team runs free for 30 days. Billing details come on day 30.
  • Cancel in Settings. Tapes and runs export as JSONL for 30 days after.
  • A workspace keeps the price it started on for 24 months.

What it does not do

Tapeback’s work ends when the pull request merges

Keep Raindrop for live traffic. Its pre-release replay, Simulations, is in research preview, and the stretch before release is where I work.

  • in scopeText agents that call toolsAlready in production, at roughly 200 conversations a week or more. Below that, read them all yourself.
  • in scopePrompt, model and tool changesAny pull request that touches one gets a replay and a blocked check.
  • in scopeRuns in your CIBudget a day to make the agent build without production secrets.
  • we don'tNo production monitoringLive tracing, issue detection and a triage agent in Slack are Raindrop’s.
  • we don'tNo verdict on last weekA bug already live replays as unchanged, and passes.
  • we don'tNo self-hostingTapes leave your network and are stored in Stockholm.

Replay last week before you ship this week

Your name and a work email get you a workspace key. If your traces already sit in Langfuse or LangSmith, last week’s tapes are in your workspace the same afternoon.

We reply within one business day.

  • ✓ Adapters today: Vercel AI SDK, LangChain, OpenAI Agents SDK, Claude Agent SDK
  • ✓ Nothing uploads until you’ve saved a redaction rule
  • ✓ The replay spends your model tokens, under your key

Your request has been received.

Expect a message from Tapeback. It goes to the address you gave.