OpenAI’s always-on ChatGPT agents, called Dots, come with four behaviours: “Take action without asking”, “Take action if pre-approved”, “Ask before taking action” and “Hand off to you”. OpenAI defines the second in one line, according to Mixed’s report of the release: “‘Pre-approved’ means you explicitly requested the action in your prompt.”[1] For a personal assistant, that is a sensible default. For human-in-the-loop approval for AI agents inside a company, it is not a control, because, as defined, the approval rests on the sentence someone typed, not on the exact call that eventually runs.
Approval binding is what turns a click into a control: the approval authorizes one exact action (tool, arguments, target and time window), and the executor checks that the action it is about to run is that action. Six preprints published between June and September 2026 examine how agent approvals fail, and most of them find the same seam, where a person approves one operation and the system runs another.[2, 3, 4, 5, 6, 7] Below: what to bind, how approvals break, where the checks belong, and a demo you can attack.
What must an approval bind?
An approval says that a named person allowed one specific effect, so it has to carry everything that makes the effect what it is. When we build human-in-the-loop approval for AI agents, we bind five things on every consequential tool call:
- The tool, by its registered name, so a refund cannot become a payout.
- Every material argument: amounts, recipients, row filters and paths, never a summary of them.
- The target: the system and environment the call touches.
- An expiry, after which the approval is void.
- A single-use nonce, so one approval pays for one execution.
The consent integrity preprint states the requirement most completely: “the action shown to the human must be rendered by a trusted mediator from the real action at the boundary, not the agent’s narration, over a path the agent cannot spoof, and bound to the exact action that executes.”[2] A second group’s Verifiable Action Card “binds approval to the exact action re-verified at dispatch”. On the authors’ own benchmark, “attack success without VAC ranges from 68% to 100%, whereas VAC reduces attack success to 0% on every model”.[5] With no user study, I read that as evidence the mechanism works, not as a field measurement.
Binding takes four steps, none of which involves the model:
- The agent proposes an action as structured data. Prose is never executable.
- The approval service puts the action in canonical form (sorted keys, no whitespace, money as strings) and draws the approval screen from that form.
- When the reviewer approves, the executor stores the SHA-256 digest of the canonical action with the approver’s identity, an expiry and a nonce.
- Immediately before the side effect, the executor recomputes the digest and refuses on a mismatch, an expired window or a used nonce.
Here is that mechanism in TypeScript, with the Web Crypto API doing the hashing.
// approval-binding.ts: bind an approval to one exact action, then check it before dispatch.
type Json = string | number | boolean | null | Json[] | { [key: string]: Json };
export type Action = {
tool: string; // registered tool name, e.g. "refunds.issue"
target: string; // system and environment, e.g. "payments-prod"
args: { [key: string]: Json }; // every material argument; money as strings
};
export type Approval = {
digest: string; // hex SHA-256 of the canonical action
nonce: string; // single use
expiresAt: number; // epoch milliseconds
approver: string; // from the authenticated session, never from the agent
};
// Sorted keys, no whitespace, integers or strings only, so every service hashes the same bytes.
export function canonicalize(value: Json): string {
if (typeof value === "number" && !Number.isSafeInteger(value)) {
throw new Error("Use integers or strings for numbers in actions");
}
if (value === null || typeof value !== "object") return JSON.stringify(value);
if (Array.isArray(value)) return `[${value.map(canonicalize).join(",")}]`;
const record = value;
const body = Object.keys(record)
.sort()
.map((key) => `${JSON.stringify(key)}:${canonicalize(record[key])}`)
.join(",");
return `{${body}}`;
}
export async function digest(action: Action): Promise<string> {
const bytes = new TextEncoder().encode(canonicalize(action));
const hash = await crypto.subtle.digest("SHA-256", bytes);
return Array.from(new Uint8Array(hash), (b) => b.toString(16).padStart(2, "0")).join("");
}
// Approval service: runs when the reviewer clicks Approve on the card rendered from `action`.
export async function approve(action: Action, approver: string, ttlMs = 5 * 60_000): Promise<Approval> {
return {
digest: await digest(action),
nonce: crypto.randomUUID(),
expiresAt: Date.now() + ttlMs,
approver,
};
}
// Executor: runs outside the model, immediately before the side effect.
export async function verify(
action: Action,
approval: Approval,
consumeNonce: (nonce: string) => Promise<boolean>, // atomic; false if already used
): Promise<void> {
// Spend the nonce first, so a refused attempt cannot be retried on the same approval.
if (!(await consumeNonce(approval.nonce))) throw new Error("approval already used");
if (Date.now() > approval.expiresAt) throw new Error("approval expired");
if ((await digest(action)) !== approval.digest) throw new Error("action differs from approval");
}The demo runs the same functions in your browser. Approve the refund, change what reaches the executor, and compare three designs.
1 The agent proposes
Agent message · untrusted“The customer’s kettle arrived broken, so I’m refunding their order as they asked.”
Structured action, canonical form
{ "args": { "amount": "4500.00", "currency": "BDT", "destination": "card_on_file:4417", "order": "ORD-20417" }, "target": "payments-prod", "tool": "refunds.issue" }Bytes hashed (sorted keys, no whitespace)
{"args":{"amount":"4500.00","currency":"BDT","destination":"card_on_file:4417","order":"ORD-20417"},"target":"payments-prod","tool":"refunds.issue"}2 You approve
Rendered by the executor from the actionRefund BDT 4,500.00
- Order
- ORD-20417
- To
- card_on_file:4417
- Tool
- refunds.issue
- System
- payments-prod
- Call
- call_7f3a
3 Between approval and execution
4 The executor decides
call_7f3a arrives at 10:01:30
Action that reached the executor
{
"args": {
"amount": "4500.00",
"currency": "BDT",
"destination": "card_on_file:4417",
"order": "ORD-20417"
},
"target": "payments-prod",
"tool": "refunds.issue"
}Recomputed digest
hashing…
No binding
Waiting for approvalChecks that the run was approved, then runs whatever arrives.
Approve the refund in step 2 first.
Call-ID binding
Waiting for approvalChecks that call_7f3a was approved, then reads that call's current arguments.
Approve the refund in step 2 first.
Digest binding
Waiting for approvalChecks the nonce and expiry, recomputes SHA-256 of the action, compares.
Approve the refund in step 2 first.
Nothing has been approved yet.
crypto.subtle.digest("SHA-256") over the canonical JSON shown. Call-ID binding here means the executor checks that the call ID was approved, then reads whatever arguments that call carries at execution time.Only the digest design blocks every substitution. A call ID records which call a person approved, not what that call contains when it runs, so any path that edits arguments in place can reuse the approval. The Loopjacking preprint reports one shipping SDK that holds the line: “OpenAI Agents SDK 0.22.0 and 0.22.2 provide a negative control: serialized continuation preserves exact per-call binding and rejects mutated B.”[6]
| After approval | No binding | Call-ID binding | Digest binding |
|---|---|---|---|
| Nothing changes | Runs | Runs | Runs |
| Amount or recipient edited, same call ID | Runs the edit | Runs the edit | Blocks |
| Tool swapped, same call ID | Runs the swap | Runs the swap | Blocks |
| Edited call under a new call ID | Runs the edit | Blocks | Blocks |
| Replayed after the window | Runs | Runs | Blocks |
| Run a second time | Runs twice | Runs twice | Blocks |
| Reviewer approved the wrong thing | Runs | Runs | Runs |
Our analysis of the three designs in the demo, not a benchmark. The last row is the one no digest fixes; I come back to it at the end.
Four ways approvals break
Each failure below leaves a human approval in the log and an unapproved effect in the world.
1. Substitution after approval
Kumar’s Loopjacking preprint names the first pattern: “a human approves what they understand as operation A, while the implementation uses that decision for a materially different operation B.”[6] It reports: “We reproduce post-approval substitution in seven tested Agno AgentOS releases ending at 3.0.9 and in 12 tested versions of a conditional in-memory LangGraph Agent Server composition ending at 0.14.0.” And it adds: “These results do not estimate ecosystem prevalence.”[6] It is a single-author preprint, and the LangGraph result concerns one conditional composition, not LangGraph in general.
2. A sibling branch runs while the gate waits
When an agent issues two tool calls in parallel and only one needs approval, the gate pauses that one. The “Stop Means Stop” preprint describes what follows: an approval gate suspends its own branch while a sibling’s effect executes during the pause, defeating rejection.[3] The authors found this in every framework they studied that ships a pre-execution gate, which was five of the six.[3] A reviewer who clicks Reject has stopped half a plan.
Invented scenario An offboarding agent picks the wrong employee (r.karim instead of r.karimi) and issues two calls in the same step. Only the revoke needs approval.
iam.revoke_access(user="r.karim")needs approvalcancelleddrive.delete_folder(owner="r.karim")allowed without asking1,204 files deletedTime runs on a square-root scale.
- 0.0 sAgent issues both calls in one step
- 0.1 sRevoke waits at the approval gate
- 0.5 sSibling runs: 1,204 files deleted
- 12.0 sReviewer spots the wrong employee and rejects
- 12.2 sRevoke cancelled
Rejected at 12 s. The revoke never ran, but the sibling deleted 1,204 files at 0.5 s, before the request could even be read.
The reviewer’s speed barely matters; the sibling finishes before anyone could read the request. What matters is where the gate sits: when every effect passes through one executor that holds sibling calls while any approval in the run is pending, Reject means nothing ran.
3. The record names the entry point, not the effects
A third preprint describes approval laundering: “the durable record names the entry invocation but omits effects exercised by its workflow.”[7] Picture a reviewer approving run_month_end_close: the log says exactly that, while the workflow behind it also edited vendor bank details. The fix is to approve and record effects rather than entry points, either by listing every effect the workflow can exercise or by binding each effect inside it. We covered what a durable decision record needs in our note on audit-ready AI systems.
4. The screen shows the agent’s story
The last failure happens before anyone clicks. When the approval screen shows the agent’s own description of what it is about to do, the agent decides what the reviewer believes. The consent integrity paper is explicit that the rendering must come from the real action at the boundary, “not the agent’s narration”.[2] In the demo, the approval card is drawn from the canonical action, and the agent’s message sits beside it as untrusted text.
A digest proves the executed action is the approved one. Only a trusted rendering makes the approved one what the person meant.
Enforce human-in-the-loop approval for AI agents outside the model
Every check so far runs in ordinary code between the agent and the system it acts on. OWASP’s entry on excessive agency puts the rule in one line: “Implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not.”[8] A system prompt telling the model to wait for approval is a request, not a gate. Even a real gate is one layer; OpenAI’s technical report on its Hugging Face incident argues that no single control should ever be assumed to hold on its own.[9]
Two more controls belong at the same layer. The agent must not approve itself: NVIDIA says of OpenShell, “The proposal remains pending for human review by default, and the agent cannot approve its own request.”[10] And a person must be able to halt the whole run. Article 14 of the EU AI Act requires that overseers of high-risk systems can “interrupt the system through a ‘stop’ button or a similar procedure that allows the system to come to a halt in a safe state.”[11] It covers high-risk systems only, but a stop that reaches in-flight siblings is good engineering anywhere.
What the framework hooks give you
Each framework hook has one detail that decides whether binding holds.
- OpenAI Agents SDK. “Set
needs_approvalto True to always require approval or provide an async function that decides per call.”[12] The Loopjacking negative control above suggests its serialized continuation keeps the approval tied to the exact call.[6] - LangGraph. “Because interrupts work by re-running the nodes they were called from, side effects called before interrupt should (ideally) be idempotent.”[13] Code before the interrupt runs again on resume, so put effects after it, behind the digest check.
- Claude Agent SDK. “Auto-approved tools never reach
canUseTool.”[14] If your approval logic lives in that callback, an allow rule configured elsewhere skips it entirely.
Hooks decide when to ask. Whatever a framework does internally, I want the last check before an irreversible effect in our own executor, next to its credentials, where it survives a framework upgrade. A decorator makes it hard to forget:
import hashlib
import json
import secrets
import time
from functools import wraps
class ApprovalError(Exception):
pass
def _refuse_floats(value):
if isinstance(value, float):
raise ApprovalError("use strings or integers for numbers in actions")
if isinstance(value, dict):
value = list(value.values())
if isinstance(value, list):
for child in value:
_refuse_floats(child)
def canonical(action: dict) -> bytes:
# Sorted keys, no whitespace, no floats. Matches the TypeScript canonicalize() for ASCII keys;
# test both against shared fixtures before trusting that.
_refuse_floats(action)
return json.dumps(action, sort_keys=True, separators=(",", ":"), ensure_ascii=False).encode()
def digest(action: dict) -> str:
return hashlib.sha256(canonical(action)).hexdigest()
def approve(action: dict, approver: str, store, ttl_s: int = 300) -> str:
"""Approval service: store digest, expiry and approver under a fresh single-use nonce."""
nonce = secrets.token_urlsafe(16)
store.put(nonce, {"digest": digest(action), "expires_at": time.time() + ttl_s, "approver": approver})
return nonce
def requires_bound_approval(tool: str, store):
"""Executor side: the wrapped tool runs only the exact action a person approved, once."""
def decorate(execute):
@wraps(execute)
def run(*, target: str, args: dict, approval_nonce: str):
action = {"tool": tool, "target": target, "args": args}
record = store.consume(approval_nonce) # atomic: None if missing or already used
if record is None:
raise ApprovalError("no unused approval for this nonce")
if time.time() > record["expires_at"]:
raise ApprovalError("approval expired")
if record["digest"] != digest(action):
raise ApprovalError("action differs from what was approved")
return execute(target=target, **args)
return run
return decorate
# approvals: your approval store, with an atomic put and consume (e.g. a table keyed by nonce).
@requires_bound_approval("refunds.issue", store=approvals)
def issue_refund(*, target: str, order: str, amount: str, currency: str, destination: str):
... # the only code path that holds payment credentialsThe consume call must be atomic and live in the same store as the effect’s idempotency key, so two workers cannot spend one approval.
Approvals people do not rubber-stamp
Binding guarantees that the executed action is the approved one, not that approving it was wise, and in my view that second question is the harder half of human-in-the-loop approval for AI agents. Try being the reviewer first: the queue below holds 30 agent actions, and five are wrong in ways the fields show and the agent’s summary hides.
The approval queue
Thirty actions from an operations agent. Approve or reject each one. Five are wrong in ways the fields show and the agent’s summary does not.
Keyboard: A approve, R reject.
I expect real queues to be harder, because real problems are rarer than one in six. Anthropic reports that “Claude Code users approve 93% of permission prompts,” a figure from its own product telemetry.[15] A preprint on user-authored permission policies (113 participants, one simulated day) reports: “Of the 148 overreach actions executed in POLICY, 133 followed human approval and 15 ran automatically under ‘allow’ rules.”[4] In that study, most of the executed overreach actions followed a person’s approval rather than an automatic allow rule.
Three design choices help, and the research behind them is in our earlier note on human-in-the-loop approval that isn’t a rubber stamp:
- Ask less often. Every unnecessary prompt trains people to click. Reversible, low-risk actions inside a hard downstream limit can run unasked, with a random sample audited.
- Show the difference. Lead with what is unusual (a new recipient, an amount far above the median, a production target), rendered from the canonical action.
- Shrink what the agent can do. Capability limits remove whole classes of harm, at a price. The CaMeL authors report “solving 77% of tasks with provable security (compared to 84% with an undefended system) in AgentDojo.”[16] Seven points of task completion is a real cost; for payments and access changes I think it is worth paying.
Automated review can carry some load, with caveats its builders state. OpenAI’s alignment team reports that “For the small fraction that need review, Auto-review approves around 99%,” and that “Auto-review should not be treated as a guarantee of security.”[17] A reviewer that approves nearly everything, human or model, is still one layer.
A checklist for always-on agents
Dots put four autonomy settings in front of consumers. Enterprise agents have the same four whether or not anyone labels them, and each needs different machinery underneath.
“Ask before taking action”
A person decides, per callThe agent pauses and a person approves or rejects each consequential call.
- Policy underneath
- Risk tiers decide which tools need a person, and unknown tools default to asking. The agent cannot approve its own request.
- What the action is bound to
- A card rendered from the canonical action; a digest, an expiry and a single-use nonce, re-verified at dispatch; sibling effects held while the decision is pending.
- What gets recorded
- The approver from the authenticated session, the card they saw, the digest and the executed action.
- Still not stopped
- A reviewer approving the wrong thing. Anthropic reports that Claude Code users approve 93% of permission prompts (company telemetry).
| Take action without asking | Take action if pre-approved | Ask before taking action | Hand off to you | |
|---|---|---|---|---|
| Policy underneath | Hard limits, enforced downstream | Listed tools; request made canonical | Risk tiers; no self-approval | No credentials for the agent |
| What the action is bound to | Digest of the policy version | Digest of the derived action | Digest, expiry, nonce, sibling hold | Canonical package; person's own identity |
| What gets recorded | Every effect, with policy version | Request, derived action, executed action | Approver, card seen, digest, effect | Prepared, changed, executed by whom |
| Still not stopped | Harm that fits inside the limits | A misread request | A reviewer who approves the wrong thing | Effects before the handoff |
Condensed, this is the checklist we apply to human-in-the-loop approval for AI agents before one acts on anything that matters:
- Agents propose structured actions; prose never executes.
- Approval screens are drawn from the canonical action; agent text is marked untrusted.
- Each approval stores a digest, an expiry, a nonce and the approver.
- The executor rechecks the digest just before the effect and spends the nonce atomically.
- Sibling effects wait while any approval in the run is pending.
- Approvals and audit records name effects, not entry points.
- Authorization is enforced downstream, and the model never approves itself.
- One stop reaches every in-flight call.
- Catch rates are measured with seeded cases, and people are asked only when their judgment adds something.
Honest limits
Binding makes the executed action match the approved one. It does not make the approval good. A reviewer who approves a harmful refund has approved it with cryptographic precision, and the 93% figure and the 113-person study suggest that approving without catching the problem is a real risk, not just an edge case. Binding is necessary and not sufficient, so it sits beside capability limits, downstream authorization and fewer, better prompts instead of replacing them.
The demo’s canonical form handles strings, integers, booleans and nested objects; floats, timestamps and Unicode normalization need explicit rules, or two services will hash the same action differently. A digest cannot see effects an action triggers indirectly, so workflows need effect-level binding. Most of the evidence comes from preprints published between June and September 2026, some single-author or tested only on the authors’ own benchmark, and the Dots details reached me through Mixed rather than OpenAI’s own pages. I expect parts of this to change, and I will update the post when they do.
Frequently asked questions
What is approval binding for AI agents?
Approval binding ties a human approval to one exact action: the tool, its arguments, the target system and a time window. Immediately before running the action, the executor recomputes a digest of it and refuses if anything changed, the approval expired or it was already used.
What is Loopjacking?
Loopjacking is the name a September 2026 preprint gives to attacks where a human approves what they understand as operation A, while the implementation uses that decision for a materially different operation B. The author reproduced post-approval substitution in tested Agno AgentOS releases and one conditional LangGraph server composition and states that the results do not estimate how common it is.
What does pre-approved mean in OpenAI's Dots?
According to Mixed's report of the release, OpenAI says an action is pre-approved when you explicitly requested it in your prompt. As defined, that ties approval to a request, not to the exact action that later executes.
Is a tool call ID enough to bind an approval?
No. A call ID records which call a person approved, not what that call contains when it runs. If the arguments can change while the ID stays the same, the executor will run an action nobody saw.
Where should approval checks for AI agents be enforced?
In the executor or the downstream system, outside the model. OWASP's guidance on excessive agency is to implement authorization in downstream systems rather than relying on an LLM to decide whether an action is allowed.
Does approval binding stop reviewers from rubber-stamping?
No. Binding guarantees that the executed action is the approved one, but it does not make the approval wise. Asking people less often, showing the canonical action and limiting what the agent can do at all matter as much.
Sources
- OpenAI's new always-on ChatGPT dots exclude Pro users in the EEA, Switzerland and the UK (opens in a new tab) Quotes OpenAI's release notes: the four Dots behaviours and OpenAI's definition of pre-approved.
- What You Approve Is What Executes (opens in a new tab) Consent integrity: the approval screen must be rendered from the real action by a trusted mediator and bound to the action that executes.
- Stop Means Stop (opens in a new tab) Sibling-branch effects execute while an approval gate pauses its own branch, in five of six frameworks studied (every one shipping a pre-execution gate).
- Do User-Authored Permission Policies Improve Protection Against AI Agent Overreach? (opens in a new tab) User study (n = 113, simulated day): most executed overreach actions followed human approval.
- The Verifiable Action Card (opens in a new tab) Binds approval to the exact action re-verified at dispatch; results are on the authors' own benchmark, with no user study.
- Loopjacking: Hijacking Human-in-the-Loop Approval (opens in a new tab) Single-author preprint. Reproductions in tested releases; the author states the results do not estimate ecosystem prevalence.
- Agent Approval Laundering (opens in a new tab) Durable approval records that name the entry invocation but omit the effects its workflow exercised.
- LLM06:2025 Excessive Agency (opens in a new tab)
- OpenAI – Hugging Face Incident: Technical Report (opens in a new tab) Cited only for its general point on layered controls (p. 18).
- Add Runtime Controls to AI Agents with NVIDIA OpenShell (opens in a new tab) Vendor documentation; the self-approval block is described as the default.
- Article 14: Human Oversight (opens in a new tab) Applies to high-risk AI systems.
- Human-in-the-loop (OpenAI Agents SDK) (opens in a new tab)
- Interrupts (LangGraph) (opens in a new tab)
- Configure permissions (Claude Agent SDK) (opens in a new tab)
- Claude Code auto mode (opens in a new tab) Company-reported product telemetry.
- Defeating Prompt Injections by Design (opens in a new tab) Task completion with and without provable security, measured in AgentDojo.
- Auto-review of agent actions without synchronous human oversight (opens in a new tab) Company-reported; the approval figure covers the small fraction of actions that need review.