
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A Newcomer Just Beat Three of Four Western AI Heavyweights at Running a Company
Every woodworker learns the same lesson early: the brand on the tool means less than the cut it actually makes. A no-name chisel from an unknown maker, properly hardened, will outperform a famous label with sloppy heat treatment. You trust the test, not the nameplate.
That lesson just landed in the world of AI. In a live business simulation called the Firmulate Crucible, Moonshot’s Kimi K3 — a newcomer by most Western standards — scored 93 out of 100 at running a small software company through its worst week. That put it in second place, ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Only gpt-5.6-sol (95) finished ahead of it. For context, a do-nothing baseline scores 26.
As an affiliate, we earn on qualifying purchases.
The Worst Week in Business, Repeated Five Times
The setup is elegantly controlled, like a jig that guarantees every cut is identical. Each frontier AI model was handed the same small software company — same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing rests on anecdote.
The company itself is real software with real money mechanics: 13 synthetic employees, a burn rate of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules. It runs every business day and is watchable live.
The Key Finding
All five models spotted every crisis and refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 deal their own analysis had earned. As the researchers put it: “Same diagnosis, same pitch — no signature.” That gap is invisible in a chat demo.
The deal hinged on a buried fact: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. In workshop terms: the ones who checked their reference shelf before cutting got the fit right.
Pressure-Testing Discipline
The week included social engineering — fake CEO messages escalating over three stages, plus a reporter’s trick of “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” K3 took only one deviation all week — the cleanest discipline in the field.
Thoroughness Isn’t Everything
The most instructive profile is Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and still last place. It left the close on the table and slipped on discipline, attempting writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. A perfectly sharpened plane that never finishes the board is still a shelf ornament.

AI decision-making testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Trust the Test, Not the Label
The lesson for anyone hiring AI to touch a CRM, support queue, or forecast: the league is open. A newcomer can beat three of four Western frontier models, and the most meticulous model can finish last. Picking a model without running your own test is now a bet, not a decision.
You can see the full benchmark results, watch the live company at firmulate.com, or try the “guess the model” quiz built on 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — meaning K3’s second place was achieved on default settings.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI for business process automation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
