Rehearse a change
before you ship it.

A software change is a hypothesis about behaviour. Replay runs it, before you have to trust it.

Proposed change
Recorded results open instantly. Run your own rehearsal live.
Live rehearsal

Investigate. Run. Compare. Verify.

Every line below is a real stage of a real rehearsal, not a progress bar for effect. Recorded results show what already happened; your own change streams these same six stages as they occur. Run current, apply change and run proposed are all carried out by the Replay Runner, deliberately not an agent: comparing two outputs is deterministic work, and a model in that path would only add cost and a failure mode.

Method

Why Replay found it

📜
Rules
Written intent: contracts, specs, a rulebook
→
🎯
Predictions
Committed in advance, quoting the rule
→
⚙
Execution
The Replay Runner. Real subprocesses, two isolated copies
→
🔍
Evidence
A verifier tries to break the result
Second opinion

A separate agent tries to break its own team's finding

The verifier is adversarial by design. It goes back into the source and the rulebook looking for a reason each finding is wrong, and reports what it could not disprove.

The verifier's evidence, line by line
No agents here

Proof it executes, computed while you read this

No model calls, nothing spent. This copies the demo application twice, applies the same fee change to one copy, and runs six hand written scenarios in each, on a payment of 100,000 rather than the 1,000 used above. The numbers below finished a moment ago, and the button replaces them with a run that happens now.

Guardrails, in code

What Replay refuses to do

Replay applies a change by finding literal text. If that text is missing, or matches more than one place, the run stops. It does not guess which one you meant, and it never reports a clean result from a change that was never applied.

services/refund_service.py: 'STANDARD_FEE_RATE' appears 3 times (lines [9, 17, 19]).
Widen the find text so it matches exactly one location, or set allow_multiple if
every occurrence really should change.

core/config.py: no occurrence of 'STANDARD_FEE_RATE = Decimal("0.999")' to replace

Both are real output from replay/sandbox.py. The first refusal is the interesting one: the refund service reads the standard fee rate in three separate places, which is exactly what breaks the refund contract above.

The rehearsal environment

AcmePay is a real, runnable system

Not a diagram. A payment service, a refund service, a modern billing service, and a legacy billing engine ported from a mainframe, all reading from a written rulebook. Replay reads it, runs it, changes it, and runs it again.

demo-repo/acmepay
View source →
ServiceScenariosResult
Honest limits

Where Replay works, and where it does not yet

✓ Executable behaviour
Functions with observable inputs and outputs.
✓ Written intent
A rulebook, a spec, documented policy to predict against.
✓ Single-call reachability
Pricing, fees, billing, tax and policy engines.
Not yet
Stateful workflows: a login, a cart, a multi-step checkout. Pointed at its own source code as a test, Replay discarded 6 of the 9 scenarios it generated because it could not execute them, and said so rather than inventing findings.

Run it on your own code

This page only rehearses against the bundled demo, on purpose: Replay executes the code it is given, so a public box accepting any repository would be executing strangers' code on this server.

git clone https://github.com/bzdmin/replay
cd replay
pip install -r requirements.txt

python scripts/check_aws.py

python scripts/rehearse.py \
  --repo /path/to/your/project \
  "Change the retry limit from 3 to 5"

You bring your own AWS credentials, and it runs entirely on your machine.

What's next

  • Scenarios with setup. Multi-step scenarios in a locked-down sandbox, so stateful behaviour becomes reachable.
  • Beyond code. Configuration, dependency upgrades and schema migrations are changes too.
  • Behavioural equivalence. Does a rewritten service still preserve the behaviour that mattered?
  • Run it on a pull request, before review, not after.
Rehearse a change before you ship it.
Source on GitHub →
AWS Agents for Humans Hackathon · Professional Agents
Strands Agents SDK · Amazon Bedrock · Apache-2.0 · ~$0.58/rehearsal