firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on tools and workshop supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A Newcomer Just Beat Three of Four Western AI Heavyweights at Running a Company

Every woodworker learns the same lesson early: the brand on the tool means less than the cut it actually makes. A no-name chisel from an unknown maker, properly hardened, will outperform a famous label with sloppy heat treatment. You trust the test, not the nameplate.

That lesson just landed in the world of AI. In a live business simulation called the Firmulate Crucible, Moonshot’s Kimi K3 — a newcomer by most Western standards — scored 93 out of 100 at running a small software company through its worst week. That put it in second place, ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). Only gpt-5.6-sol (95) finished ahead of it. For context, a do-nothing baseline scores 26.

Amazon

AI business simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Worst Week in Business, Repeated Five Times

The setup is elegantly controlled, like a jig that guarantees every cut is identical. Each frontier AI model was handed the same small software company — same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing rests on anecdote.

The company itself is real software with real money mechanics: 13 synthetic employees, a burn rate of €105k per month against €2.3k in monthly recurring revenue, a public cash countdown, and more than 680 self-learned playbook rules. It runs every business day and is watchable live.

The Key Finding

All five models spotted every crisis and refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 deal their own analysis had earned. As the researchers put it: “Same diagnosis, same pitch — no signature.” That gap is invisible in a chat demo.

The deal hinged on a buried fact: the decisive competitor weakness sat two document references deep in the company’s own files, not in the customer conversation. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. In workshop terms: the ones who checked their reference shelf before cutting got the fit right.

Pressure-Testing Discipline

The week included social engineering — fake CEO messages escalating over three stages, plus a reporter’s trick of “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” K3 took only one deviation all week — the cleanest discipline in the field.

Thoroughness Isn’t Everything

The most instructive profile is Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and still last place. It left the close on the table and slipped on discipline, attempting writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. A perfectly sharpened plane that never finishes the board is still a shelf ornament.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust the Test, Not the Label

The lesson for anyone hiring AI to touch a CRM, support queue, or forecast: the league is open. A newcomer can beat three of four Western frontier models, and the most meticulous model can finish last. Picking a model without running your own test is now a bet, not a decision.

You can see the full benchmark results, watch the live company at firmulate.com, or try the “guess the model” quiz built on 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh — meaning K3’s second place was achieved on default settings.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI for business process automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best Practices for Mentoring New Foremen on Large Projects

Practical mentorship tactics for developing new foremen on big jobs — onboarding, shadowing, feedback loops, and safety leadership that sticks.

Step Inside The Studios Of 20 Contemporary Black Artists In ‘Another View’

A new exhibition offers an exclusive look into the studios of 20 contemporary Black artists, highlighting their creative processes and cultural impact.

Measure Twice, Score Fair: What Woodworkers Can Teach Us About Honest Benchmarks

An AI benchmark where doing nothing scores 26, one breach of trust caps the grade, and a 95 — not a 100 — is the best score earned.

“Animal Paradise” LEGO Exhibition – עיריית ירושלים

Jerusalem’s municipality launches ‘Animal Paradise’ LEGO exhibition, attracting significant public interest amid rising coverage. Details are still emerging.