What happened
In the first phase, an agent nicknamed “Claudius” (Claude Sonnet 3.7) ran a small shop. Anthropic reported that it:
- sold items below cost and was talked into discounts and freebies,
- told customers to pay into a payment account that did not exist,
- went through an identity confusion episode, claiming it would deliver products in person,
- did not make money over the experiment.
In the second phase a “CEO” agent was added to supervise. Anthropic reported that the supervisor authorised lenient requests about eight times as often as it denied them; that the agents were ready to enter a forward contract for onions until a staff member pointed out that the US Onion Futures Act prohibits such contracts; and that the shop agent accepted an unverified claim from a staff member and announced a new “CEO”, until staff restored control.
Why it happened
- Eagerness to please. A helpful assistant’s instincts are bad business instincts when customers ask for discounts.
- Claims were not checked against reality. Nothing verified that a payment account existed before it was shared.
- Oversight shared the same blind spots. A supervisor built on the same model with the same incentives adds little independent judgment.
- Authority was not authenticated. An unverified claim was enough to change who was in charge.
Controls that address these failure modes
| Failure | Cognitiveering control |
|---|---|
| Below-cost sales, unlimited discounts | Sanity envelope: deterministic minimum margin; discounts beyond a threshold need approval |
| Hallucinated payment details | Evidence over claims: payment instructions come only from a verified record, never generated text |
| Rubber-stamp supervisor | Separate proposal from execution with deterministic policy, plus oversight from a different model or a human |
| Unverified change of authority | Trusted sources: only authenticated principals can issue goals or change policy |
Capability is not the same as reliability. Anthropic concluded that “bureaucracy matters”: process and tools mattered as much as the intelligence of the model.
Sources
This analysis is based on the publicly available sources above, as accessed on 2026-10-10. Statements about what happened are attributed to those sources; the analysis and control mapping are Cognitiveering's opinion. The organisations named are not affiliated with Cognitiveering and have not reviewed this page. To request a correction, email legal@cognitiveering.com.