Tapeback
Replays last week's real conversations against your next agent change
Track 01/07
Tapeback serves the engineer who owns an agent that talks to paying users, such as a support or booking agent.
That engineer changes its prompt, model or tools several times a month, checks a few dozen saved cases, and ships. A bad reply still returns HTTP 200, so nothing alerts and the change is tested on customers.
Track 02/07
Tapeback was founded in March 2022, before LLM agents were something a team put in front of paying users.
The first product replayed logged requests at a changed service and compared responses by structure, which is all a rule can decide. Once the likeliest change became a prompt, the difference was text, and the second product cannot work without a model.
Track 03/07
Tapeback records production conversations as tapes, redacted in the customer's own process before upload.
When a pull request touches a prompt, a model setting or a tool, a CLI in the customer's CI replays a sample of the past week's tapes against the changed agent. Changed steps come back in groups, and the check stays blocked until a named person marks each group.
Track 04/07
Rules can compare the structure of two conversations but are silent on their text, so Tapeback runs two models.
Whether a reworded step means the same thing is a judgment about meaning, and we plan for about 85% of changed steps to get no verdict from a rule. A pair classifier gives each such step one of six labels and a confidence; an embedding model groups steps that changed alike.
Track 05/07
A classifier label sorts a Tapeback report but never clears a change.
Below 0.80 confidence no label is shown, and an engineer labels the step by hand. Every group stays Unmarked until a person marks it Expected or Regression. Each mark leaves a labeled pair, and we train only on pairs from opted-in workspaces, never on tapes.
Track 06/07
Each pull request sets off a burst of embedding and labeling on Tapeback's GPUs.
A typical sample means about 11,000 step pairs in a few minutes while a CI check waits, then nothing until the next change. Tapes hold what end users typed, so both models run on cloud GPU capacity we control in Stockholm, where Tapeback is registered, and tapes stay on cloud infrastructure in Stockholm.
Track 07/07
A smaller classifier of our own replaces Tapeback's prompted 8B pair labeler by March 2027.
In the fourth quarter of 2026 that labeler, an open-weight 8B model, shares cloud GPU capacity we control in Stockholm with the embedding model. Its replacement is fine-tuned there on opted-in pairs, and by September 2027 retrieval over a workspace's past marks suggests marks for familiar groups, for a person to confirm.
Linnea Sjöstrand
Founder
Owned the refund service at an e-commerce software company and shipped a change that passed every saved test, then refunded against the wrong order for two days before anyone saw it.
Filmon Tewolde
CTO
Spent four years training models that decide whether two sentences mean the same thing for a translation quality team; the pair labeler is that problem with tool calls in it.
Mikaela Grönroos
CPO
Ran release engineering for a payments API in Helsinki where no change shipped without a replay of recorded bank traffic, and could not find the same habit anywhere in agent teams.