P01
Decision Record
Problem. Decisions are forgotten, rationalised after the fact, and never learned from.
- For people
- Before deciding, write down the intent, options, expected outcome, and your confidence. Review it when the outcome is known.
- For agents
- Log intent → policy check → approval → action → evidence → outcome as one durable record per consequential action.
Failure mode
Logging only the final answer, so nobody can reconstruct why it happened.
P02
Premortem
Problem. Plans look safe because nobody is asked how they could fail.
- For people
- Imagine it is a year from now and the plan failed. List the most likely reasons, then mitigate the top three.
- For agents
- Before an irreversible action, have the planner list failure scenarios and check each against deterministic policy.
Failure mode
Running the premortem with the same blind spots, e.g. the same model judging its own plan.
P03
Separate Proposal from Execution
Problem. Whoever proposes an action also gets to approve it.
- For people
- Use a cooling-off rule: big decisions are proposed one day and confirmed the next, ideally by a second person.
- For agents
- The model proposes; an independent gateway authorises. Prompts such as "do not spend money" are not a control.
Failure mode
A supervisor that rubber-stamps because it shares the proposer’s incentives or model.
P04
Hash-Bound Approval
Problem. An approval is given for one thing and used for something slightly different.
- For people
- Approve the exact document version, not "the contract". Any edit means approving again.
- For agents
- Bind each human approval to a hash of the exact action payload, with an expiry and single use; re-check it at execution time.
Failure mode
Approving "send the email", then the agent edits the recipient list.
P05
Evidence over Claims
Problem. "Done" is reported without proof that anything actually changed.
- For people
- Close a task only when you can point to the artifact: the receipt, the merged PR, the signed form.
- For agents
- An HTTP 200 is not success. Read back the external state and store the evidence before marking the goal complete.
Failure mode
An agent confidently reports a payment, order, or account that never existed.
P06
Sanity Envelope
Problem. A single typo or misreading produces an absurd outcome.
- For people
- Define normal ranges in advance (budget, quantity, time) and pause when you are outside them.
- For agents
- Deterministic bounds on quantities and amounts relative to history; anything outside the envelope needs approval.
Failure mode
Ordering thousands of units when the usual order is dozens.
P07
Calibration Loop
Problem. Confidence does not track accuracy, and nobody measures it.
- For people
- Attach a probability to predictions and score yourself over time.
- For agents
- Compare stated confidence against verified outcomes; route low-calibration task types to stronger models or humans.
Failure mode
Measuring activity (messages sent, hours run) instead of verified outcomes.
P08
Trusted Sources
Problem. Instructions and facts are accepted from anyone who sounds authoritative.
- For people
- Verify who is asking through a second channel before acting on unusual requests.
- For agents
- Tag every memory and instruction with its source and trust level; untrusted content can never grant permissions.
Failure mode
An agent accepts an unverified claim of authority, or a prompt injected into a web page.