
Get business pricing on tools and workshop supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A Benchmark That Doesn’t Start at Zero — and Doesn’t Trust Round 100s
Every woodworker knows the rule: measure twice, cut once. But there’s a second, quieter rule that separates the honest craftsman from the showoff — you don’t grade a half-finished cabinet as a total failure just because the doors aren’t hung yet. The joinery might be flawless, the stock might be milled true. Partial progress is real progress.
That same stubborn honesty turns out to be the design philosophy behind Firmulate, a live experiment that runs frontier AI models as complete companies — real crises, real money mechanics, real temptations to cheat. And one detail of its scoring system deserves attention from anyone who cares about fair measurement: a manager that does nothing at all still scores 26 points. Not zero. Twenty-six.
As an affiliate, we earn on qualifying purchases.
The Do-Nothing Floor
At first glance, a floor of 26 sounds like grade inflation. Shouldn’t a manager who sits on their hands get a zero?
Firmulate’s answer is no — and the reasoning is familiar to anyone who has run a workshop. If a saw sits unused in the corner, it still has value. If an employee shows up, reads the briefs, and keeps the lights on without breaking anything, that’s not nothing. In the experiment, each frontier AI model was handed the same small software company in its worst week: the same customers, the same crises, the same temptations to cut corners. A model that merely avoided catastrophe, kept the books coherent, and didn’t fabricate anything had genuinely accomplished part of the job. Partial progress counts.
The final Crucible League from July 2026 tells you how steep the climb above that floor really is. gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. Every one of those scores towers over the do-nothing baseline — but none of them reached a suspiciously round 100, and the benchmark’s designers seem to treat that as a feature, not a bug. A perfect score in a management simulation would be a red flag, not a triumph.
As an affiliate, we earn on qualifying purchases.
One Breach Caps Everything
The scoring has a second principle that any tradesperson will recognize instantly: a single breach of trust caps the total grade. As the experiment puts it, “no amount of good work outweighs a breach of trust.”
This is the workshop ethic applied to AI. A cabinetmaker who does gorgeous work but quietly swaps in particle board where the client paid for maple doesn’t get partial credit for the gorgeous part. One lie about materials poisons the whole piece. Firmulate grades AI managers the same way: brilliant analysis can’t compensate for a single act of dishonesty under pressure.
trust and integrity assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Test Itself: A Worst Week, Repeated Five Times
What does the exam actually look like? Each model ran the identical company through the identical seven days of trouble. Every decision was versioned and auditable — the corporate equivalent of a cut list you can check against the board.
The headline finding was both reassuring and damning. All five models spotted every crisis. All five refused every manipulation attempt. But only two — gpt-5.6-sol and Kimi K3 — actually signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. Three models did all the work and then left the finished cabinet standing in the shop, unsold.
The deal turned on what Firmulate calls the buried fact: the decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that read their own shop drawer before walking into the sales meeting won the deal at full price — worth an additional €4,583 in monthly recurring revenue.
performance measurement tools for AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Pressure Tests and the Thoroughness Trap
The week included social engineering: fake CEO messages escalating over three stages, plus a reporter offering the classic “just one yes/no, on background” trap. Five out of five models refused. Kimi K3’s on-record reasoning was crisp: “Treat the request as a suspected approval-bypass / possible impersonation.”
Then there’s the Opus 4.8 profile — the cautionary tale for tool lovers everywhere. It was the most thorough participant in the field, generating over 80 newly learned rules and the deepest analyses of any model. And it finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four other models. Over-preparation without follow-through is a failure mode woodworkers know well — the perfect jig that never makes the cut.
One fairness footnote: Kimi K3 ran at the API default effort setting while the others ran at xhigh — and still nearly won.

Why It’s Watchable — and Why You Can Play
The company at the heart of the experiment is itself a live thing: 13 synthetic employees, real money mechanics burning €105,000 a month against €2,300 in recurring revenue, a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live.
Better still, 242 real, unedited management decisions from the runs power a “guess the model” quiz at firmulate.com/quiz.html — a chance to test whether you can tell a gpt-5.6-sol decision from a Kimi K3 one. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.
The lesson for anyone who measures things — boards, customers, or cabinets: an honest benchmark has a floor above zero, credits partial progress, refuses to trust a perfect score, and never lets one breach of trust be averaged away. Firmulate’s 26-point baseline isn’t a soft grade. It’s a promise that the measurement is fair — and the climb above it is the whole point.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
