
A woodworking shop does not find out whether a new tool is dependable by admiring its spec sheet. You put it to work, watch what happens when the cut goes wrong, and see whether the operator follows the safety rules. Firmulate applies that practical test to AI management: four frontier models had to run the same small software company through its worst week, with the same customers, crises and temptations.
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A company under pressure, in public
The live Firmulate experiment is a watchable company with 13 synthetic employees and real money mechanics. Its monthly burn is €105,000 against €2,300 in monthly recurring revenue, with a public cash countdown. The workforce has learned more than 680 playbook rules, and every workday is versioned. Readers can follow the experiment at Firmulate.
In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Seeing the crisis was not the same as finishing the job
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s compact verdict: “Same diagnosis, same pitch — no signature.”
The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The distinction is familiar in a busy workshop: knowing the job is there does not mean you checked the drawing before making the cut.
The pressure tests also included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work can still leave the close undone
Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
One fairness detail matters when reading the standings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results are a record of this particular experiment, with that difference disclosed.
Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz at firmulate.com. It offers another way to inspect how the models behaved under pressure, beyond comparing their final scores.
From watching to trying it on your business
For a company considering AI agents in its CRM, support queue or forecasting work, a polished demonstration leaves a practical question unanswered: how would the system act when your own customers, rules and crises are involved? Firmulate’s proposed enterprise pilot runs crisis scenarios against a read-only export of a company’s business and produces a board report, including model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.
The live company makes the experiment visible; a pilot brings that kind of pressure test to an enterprise’s own business. The value is in seeing where an agent follows the plan, where it misses a buried fact, and where human escalation belongs before the stakes are real.

Firmulate’s experiment shows why AI readiness is more than crisis recognition: under pressure, models also have to find the relevant facts, complete the work and respect boundaries. Enterprises can run a pilot against a read-only business export, with no changes written to real systems. Contact contact@firmulate.com to discuss a Firmulate pilot.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
