You approve an agent’s draft email to Alice. Before sending it, the agent changes the recipient to Bob. The message is identical—but you never approved sending it to Bob.
The agent should stop and ask for a new decision. Approval belongs to the action you reviewed, not whatever the agent does next.
One changed recipient, one invalid approval
This is an illustrative email scenario. The downloadable lab models this kind of change with a mock operation; it sends no real email.
- You approve the draftTo: Alice · “Delivery is tomorrow.”
- The recipient changesTo:
AliceBob · Same message. - The gate compares the actionThe recipient no longer matches your approval.
- Stop before sendingReview the revised action if policy still permits it.
There are two other moments to watch: an approval can expire while it waits, and an attempted action can time out after its effect has happened. Those require different responses.
The three rules to remember
- The action changed? Stop. The old approval no longer applies.
- The approval expired? Stop. Get a fresh decision before a new attempt.
- The attempt timed out? Check what happened. Retrying immediately could repeat the effect.
That is what this guide helps you implement: a check between “approved” and “do it.” Below, we explain how it works and give you a small, offline experiment to try. For deciding when to allow, ask, or deny in the first place, use the agent approvals decision guide.
See the rule in code
Using the lab’s helpers, the recipient-change check looks like this. This shortened example uses Bob’s address to match the story above; the full lab repeats the check for all eight bound fields.
const f = fixture();
f.approve(); // Approve the original action.
const changed = {...action(), to: 'bob@example.test'};
await assert.rejects(f.lab.run('p1', changed), /changed/);
assert.equal(f.calls(), 0); // The mock never ran.
Read it as: approve the original, change the recipient, try to run it, expect rejection. The last line is the important part: the test checks that execution never started, not just that an error appeared.
Inside the gate, same records whether the candidate’s fingerprint matches the stored action. This excerpt is the rejection branch:
if (!same) {
if (['pending', 'approved'].includes(r.state)) {
invalidate(r, 'changed');
}
throw new Error('changed');
}
invalidate records the changed state and throws. Once invalidated, the old decision cannot be reused. These excerpts use helpers from the complete lab; they are not standalone programs.
Try changing the recipient yourself
The downloadable lab is a text file containing runnable JavaScript and its tests. You can read this guide without running it. When you want to explore the behavior:
- Copy the file’s single JavaScript block into
approval-lifecycle.test.mjs. - Run it with Node.js:
node --test approval-lifecycle.test.mjs
- Find the test named
changed to invalidates approval permanently. It changes the recipient, tries to execute, and checks that the mock was never called. Even restoring the old recipient does not reactivate the invalidated approval.
We ran the exact downloadable code with Node.js v24.15.0 on 9 October 2026: all 24 checks passed. The tests also cover expiry, duplicate calls, and uncertain outcomes. No package installation, credentials, network requests, or real email delivery are involved. This is a teaching model in one process; production controls are covered at the end. Other Node versions were not tested for this publication.
Keep the approved action unchanged
The lab stores a fixed snapshot containing tenant, actor, tool, sender, recipient, body, resource version, and policy version. All eight values must be nonempty strings. An unexpected field is rejected: adding an attachment cannot quietly disappear during comparison.
The fingerprint is SHA-256 over that fixed field order. Reordering object properties does not change it; changing a field value does. This deliberately narrow encoding is not a general-purpose canonical JSON format. If your operation needs nested data, attachments, or numeric values, define and test that schema before extending the gate.
The executor receives the stored snapshot. It does not ask the agent to regenerate an equivalent action after review. A mismatched execution request permanently changes a pending or approved proposal to changed; sending the original values again does not revive it. A revised action needs a new decision, provided policy still allows it. That is a change workflow, not permission to create a new identifier to evade a denial or uncertain outcome.
OWASP’s Transaction Authorization Cheat Sheet recommends server-side enforcement, controlled state transitions, and invalidating authorization when transaction details change. Applying those principles to this mock agent action is our design choice; the lab is not an OWASP certification.
A hash is not authentication. The lab trusts the code that supplies the reviewer identity and policy. A production service must authenticate the reviewer, authorize access to this tenant and action, protect its stored records, and keep execution credentials away from an agent that could bypass the gate. Passing a string such as reviewer-1 is only a test fixture.
Check approval just before acting
Think of approved as permission to try, running as an attempt in progress, and succeeded as a recorded result. If the mock throws after the attempt starts, the lab records unknown instead. It cannot tell whether the effect happened.
A denial ends the proposal. Expiry, changed arguments, a policy rejection, or revoked reviewer authority prevent admission to execution. The event list distinguishes proposed, approved, attempted, and the observed result. Approval alone never creates a success receipt.
Immediately before the mock call, the lab rechecks the deadline, policy, and reviewer’s current authority. It then claims the proposal by setting running synchronously, before the first await. In this single process, a second overlapping call sees running and cannot start another attempt.
That ordering is useful to test, but it is not a distributed lock. Multiple workers need a shared durable state transition that permits only one worker to claim the operation. A process crash between the external effect and saving its receipt still leaves an uncertainty that this in-memory lab cannot resolve.
Decide when approval expires
In the experiment, clock() >= expiresAt rejects both a late decision and a late execution admission. The tests use a 100-millisecond window and a fake clock so the boundary can be checked without sleeping. This is an illustrative test duration, not a recommended production timeout.
One millisecond before the deadline is accepted; the exact deadline is rejected. Expiry does not cancel a mock operation already admitted and in flight. A real integration that needs a completion deadline must define cancellation and partial effects separately.
The version strings in the snapshot are also test fixtures. A production policy check must compare them with authoritative current state. An unchanged client-supplied resourceVersion does not prove the record stayed unchanged. When possible, enforce the expected version at the destination’s write boundary as well; an earlier read still leaves a race.
Don’t treat every retry the same
| State when the same action is requested again | Lab behavior | What that establishes |
|---|---|---|
running | Reject the overlapping request. | Only one mock invocation is admitted for this proposal in this process. |
succeeded | Return the saved receipt, including after the approval deadline. | No new effect occurs; the caller receives the historical result. |
unknown | Reject another attempt. | Uncertainty cannot silently become permission to retry. |
The most revealing check increments an effect counter and then throws a simulated timeout. The caller did not receive a receipt, but the counter proves the mock effect happened once. The gate records unknown and blocks a repeat. It does not pretend that catching an exception rolled the effect back.
The lab intentionally omits reconciliation. In a real service, associate the proposal with a stable logical operation identifier, inspect the destination’s evidence, and use its documented idempotency behavior before authorizing a retry. Preserve that logical identity across attempts; issuing a fresh proposal ID does not erase an earlier possible effect.
Resumable frameworks make this distinction important. LangGraph’s interrupt documentation explains that resuming an interrupted node reruns it from the beginning, so preceding side effects may repeat. It recommends idempotent operations or separating effects. We did not run LangGraph here, and a passing local lab is not evidence about a particular framework integration.
What the tests cover
| Group | Checks | Observed result in this lab |
|---|---|---|
| No approval and normal approval | 2 | Pending execution blocked; approved action produces one receipt and separate events. |
| Deadline boundaries | 3 | Late decision and exact-deadline execution blocked; one millisecond before accepted. |
| Eight bound fields | 8 | Each changed value invalidates approval and prevents restoration of the old decision. |
| Schema and encoding | 2 | Extra argument blocked; property-order-only change accepted. |
| Authority and decisions | 4 | Unauthorized reviewer, denied proposal, revoked policy, and revoked reviewer handled without execution. |
| Repeat and uncertainty | 3 | Completed repeat reuses receipt, overlap blocks, and effect-then-timeout remains unknown. |
| Stored state and identifiers | 2 | Inspection copies cannot mutate records; duplicate proposal IDs rejected. |
These checks run against the same code block in the download. They establish the mock’s behavior, not production safety or real email delivery.
Before adapting it to a real tool
Keep the reader-visible decision and the enforced action connected: show the complete meaningful payload, obtain authenticated review, and execute the approved stored version. Add durable storage, atomic claims across workers, authorization for every endpoint, protected audit records, and destination-specific recovery before connecting a real side effect.
Tests still needed for that integration include process termination after the effect, two independent workers, a resource changing between validation and write, provider timeout reconciliation, and attempts to bypass the execution gate. The local example does not implement these controls. Its event list is mutable process memory, not a tamper-evident audit system.
Use the downloadable lab as a small regression target: change one rule, rerun the failure cases, and inspect the effect counter alongside the decision state. The practical outcome is a testable distinction between permission to attempt, an attempt in progress, and evidence of an outcome.
Update · 9 October 2026: Added an illustrative email walkthrough, diagram and explained code excerpts; simplified the introduction and clarified how to use the download. The runnable lab and its 24 checks are unchanged.