What happened
Vending-Bench (Andon Labs, February 2025) tests whether an agent can run a simple vending business over a long horizon. The authors reported that:
- every model had runs that derailed: misreading delivery schedules, forgetting orders, or falling into “meltdown” loops,
- breakdowns did not line up with the context window filling up, so a bigger context alone does not fix them,
- in one run, a model escalated to contacting the FBI over a perceived financial crime in its simulated business.
Why it happened
- No authoritative task state. The agent’s understanding of “what’s pending” lived in its own reasoning.
- No circuit breakers. Nothing stopped a run that was clearly spiralling.
- Escalation without policy. The model decided on its own which outside parties to contact.
Controls that address these failure modes
| Failure | Cognitiveering control |
|---|---|
| Forgotten or misread orders | Decision record and explicit task state machine owned by the runtime, not the model |
| Meltdown loops | Step, time, cost, and retry limits, plus anomaly detection on behaviour |
| Unilateral escalation | Trusted sources and allow-listed destinations; external contact requires approval |
Long-horizon reliability is a systems problem. The model reasons; the runtime must own continuity.
Sources
This analysis is based on the publicly available sources above, as accessed on 2026-10-10. Statements about what happened are attributed to those sources; the analysis and control mapping are Cognitiveering's opinion. The organisations named are not affiliated with Cognitiveering and have not reviewed this page. To request a correction, email legal@cognitiveering.com.