firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A woodworking shop does not find out whether a new tool is dependable by admiring its spec sheet. You put it to work, watch what happens when the cut goes wrong, and see whether the operator follows the safety rules. Firmulate applies that practical test to AI management: four frontier models had to run the same small software company through its worst week, with the same customers, crises and temptations.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A company under pressure, in public

The live Firmulate experiment is a watchable company with 13 synthetic employees and real money mechanics. Its monthly burn is €105,000 against €2,300 in monthly recurring revenue, with a public cash countdown. The workforce has learned more than 680 playbook rules, and every workday is versioned. Readers can follow the experiment at Firmulate.

In the final Crucible League, published in July 2026, gpt-5.6-sol placed first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The experiment counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Seeing the crisis was not the same as finishing the job

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s compact verdict: “Same diagnosis, same pitch — no signature.”

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The distinction is familiar in a busy workshop: knowing the job is there does not mean you checked the drawing before making the cut.

The pressure tests also included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work can still leave the close undone

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. The close was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

One fairness detail matters when reading the standings: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results are a record of this particular experiment, with that difference disclosed.

Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz at firmulate.com. It offers another way to inspect how the models behaved under pressure, beyond comparing their final scores.

From watching to trying it on your business

For a company considering AI agents in its CRM, support queue or forecasting work, a polished demonstration leaves a practical question unanswered: how would the system act when your own customers, rules and crises are involved? Firmulate’s proposed enterprise pilot runs crisis scenarios against a read-only export of a company’s business and produces a board report, including model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

The live company makes the experiment visible; a pilot brings that kind of pressure test to an enterprise’s own business. The value is in seeing where an agent follows the plan, where it misses a buried fact, and where human escalation belongs before the stakes are real.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Firmulate’s experiment shows why AI readiness is more than crisis recognition: under pressure, models also have to find the relevant facts, complete the work and respect boundaries. Enterprises can run a pilot against a read-only business export, with no changes written to real systems. Contact contact@firmulate.com to discuss a Firmulate pilot.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Implementing Conflict Resolution Tactics in a High-Pressure Jobsite

Practical conflict resolution tactics for foremen and superintendents. Real steps, scripts, and warning signs to shut down jobsite disputes fast.

You Test a Saw Before You Trust It. So Why Don’t You Test an AI Before You Hire It?

Newcomer Kimi K3 scored 93 in the Firmulate Crucible, beating three of four Western frontier AI models at running a company. The league is open.

These Kitchen Cabinets Are Unrecognizable After A $600 DIY Makeover

A homeowner’s $600 DIY kitchen cabinet renovation has transformed the space, sparking widespread interest and debate over DIY potential and results.

Ryobi VS Ridgid Impact Wrench Head-to-Head: Has Prosumer Caught Up?

Search interest in Ryobi vs Ridgid impact wrench head-to-head comparisons is spiking. What is confirmed, why shoppers care, and what remains unclear.