You cannot review a prompt diff
A short introduction to Rohan Prashanth's full essay on the Space AI blog.
One edited sentence in an agent's instructions can change what it does with your money, and no code review will catch it. Our co-founder Rohan Prashanth wrote about the day this happened to us, and what we built so it can't happen again.
What broke
We edited one section of an agent's prompt to make it more decisive. The change was short and read cleanly in review. The next day, operators started seeing proposals they had already been turning down. Nothing crashed. Every test passed. The only signal was that people trusted fewer suggestions.
That is the risk with AI in a real business: it fails quietly, through judgement, not errors.
What we built
Every agent action already goes through an approval queue, so we have a record of what people accept and refuse. We turned that record into a test. Before any prompt change ships, it is replayed against recent real situations in read-only mode. If the new version starts reaching for actions people usually reject, the change is blocked.
Rohan's essay covers the details, including what we got wrong along the way:
- Rebuilding past situations silently drops context, so we now store the exact inputs.
- Comparing wording is useless; comparing the type of action proposed is what maps to business risk.
- If the model in use changes during the test, you can't tell the prompt's effect from the model's.
It also says plainly what the gate does not do: approval is a stand-in for good judgement, not proof of it.
Why we're sharing it
Most businesses adopting AI will never write a prompt. But every one of them should ask whoever builds their agents: when you change how this behaves, how do you know it still makes decisions we'd agree with?
If the answer is "we read it and it looked fine," that is not good enough.
Read the full essay: You Cannot Review a Prompt Diff, by Rohan Prashanth, 17 September 2026.
Originally published on the Space AI blog.