
Imagine an AI handling a beauty brand’s customer complaints, supplier delays and pricing decisions during a rough week. A polished answer in a demo may look reassuring. The harder question is whether the system can spot a risk, stick to the rules and follow through when a real sale is on the line. Firmulate is testing that question by putting AI models in charge of a company under pressure.
Get beauty and skincare favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate’s Crucible League ran frontier models through the same small software company and its worst week: identical customers, crises and temptations. Every decision was versioned and auditable. In the final results from July 2026, gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. The league’s rule is plain: partial progress counts, but one breach of trust caps the total. “No amount of good work outweighs a breach of trust.”
Spotting trouble is only the start
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap between recognizing the right move and carrying it through is the story: same diagnosis, same pitch — no signature.
The deciding clue was easy to miss. A competitor’s weakness sat two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. That result makes the exercise relevant well beyond software: in beauty and personal care, important context may live in product, customer or operating records, while the urgent problem arrives somewhere else.
Trust under pressure
The models faced fake CEO messages that escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Refusal alone did not guarantee strong performance. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of the same problem appeared in all four models. For a company considering AI in customer care, marketing or operations, following the right boundary matters alongside knowing where it is.
Watch, then test against your own business
The live company has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record for every workday. The experiment is real and watchable at firmulate.com. A quiz built from 242 real, unedited management decisions invites readers to guess which model made each choice.
There is a fairness caveat in the league comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results are a specific experiment, not a promise that an AI will behave identically in another company or setting.
For an enterprise, the next step is to try the wargame on its own business. Firmulate says a pilot can use a read-only export, run crisis scenarios against that company’s context and produce a board report with a model ranking and weak points in its playbooks. Nothing writes back to real systems. The pilot takes the experiment from watching a synthetic company to examining how models might handle yours.

Put the playbooks to the test
A model can identify a crisis and still miss the decision that matters. Firmulate’s pilot lets enterprises wargame scenarios against a read-only export of their own business, with no writes to real systems. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
