
Wellness depends partly on whether the systems around us can handle pressure without passing the damage along. As AI moves into decisions that affect customers and workers, companies have a new question to face: what will these agents do when the week goes wrong?
Get health and wellness essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate puts that question to a live, watchable experiment. Its public brand runs AI models as companies under pressure, then measures what they decide and whether they follow through. The next step is taking that kind of rehearsal into an enterprise’s own business.
When the crisis is familiar, follow-through matters
In the final Crucible League, published in July 2026, each frontier model faced the same small software company, customers, crises and temptations. Every decision was versioned and auditable. The standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26.
The headline finding was less about spotting trouble than acting on it. Every model identified every crisis and refused every manipulation attempt. But only two signed a €55,000 deal that their own analysis had earned. The researchers summed up the gap: “Same diagnosis, same pitch — no signature.” In a company, recognizing the right move is not the same as making it.
The detail was buried in the company’s own files
The decisive competitor weakness did not appear in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. That detail turns the exercise into a test of whether an AI can connect evidence inside a business to a decision under pressure.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a useful caution against equating thoroughness with performance. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal unsigned and discipline slipped: it attempted to write into a locked department instead of escalating. A weaker version of the same issue appeared across all four models. For fairness, K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
From watching to rehearsing your own business
Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and a versioned record for every workday. The live experiment can be watched at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made them.
For a business considering AI agents in customer support, sales or forecasting, the more personal question is what happens when the agent meets that company’s own constraints, playbooks and pressure points. Firmulate’s proposed enterprise pilot uses a read-only export to create a digital twin, then runs crisis scenarios against it. The output is a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.
That boundary matters: a rehearsal can expose gaps without handing an experiment the keys to live operations. The results could help leaders see where a model recognizes a problem but fails to act, or where internal guidance leaves room for a costly mistake.

Test the response before the real crisis
Firmulate’s league suggests that models can spot crises and resist manipulation while still missing a deal or mishandling authority. A company-specific wargame offers a way to examine those behaviors against its own business before agents are trusted with consequential work.
To explore a pilot using a read-only export and scenarios based on your company, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
