
We all know someone who coasts. Shows up, does the bare minimum, signs nothing, breaks nothing, offends no one — and somehow still gets a decent performance review. It feels like grade inflation. But what if doing nothing actually is worth something — just not very much?
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
That’s the quietly radical idea behind Firmulate’s benchmark, an experiment that ran four frontier AI models as managers of the same small software company through its worst week. When the organizers scored a “do-nothing” baseline — a manager who simply sits on its hands — it got 26 points out of 100. Not zero. And the reasoning behind that number says a lot about how we should measure people, and machines, honestly.
First, the experiment
Firmulate gave each AI model the identical job: run a small software company through a brutal stretch — same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing could be quietly rewritten after the fact. The final July 2026 league table: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the floor is 26, not 0
Here’s the part that trips people up. A manager who does absolutely nothing still avoids catastrophic mistakes. The company doesn’t get defrauded, no customer gets lied to, no locked department gets broken into. Partial progress counts: if you spot the crisis but don’t resolve it, that’s worth more than missing it entirely. If you diagnose the problem correctly but never close the deal, that’s worth more than a wrong diagnosis.
So the do-nothing baseline earns 26 points — a recognition that restraint has value, but nowhere near enough to win. It’s the difference between grading on a curve that flatters everyone and grading that tells you exactly how much of the job actually got done.
business decision analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why nobody scores a suspicious 100
The same honesty applies at the top. The benchmark has a built-in distrust of round numbers, and its structure explains why: a single breach of trust caps the total grade. As the rules put it, “no amount of good work outweighs a breach of trust.” One lie to a customer, one forged approval — and no pile of brilliant decisions can buy back a perfect score. That’s why the winner landed at 95 rather than 100. Real management, it turns out, always leaves something on the table.
employee performance evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What actually separated the winners
The headline finding was humbling for the machines: all five models spotted every crisis and refused every manipulation attempt. Where they diverged was finishing. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The deal-clinchers won because they found a decisive competitor weakness buried two document references deep in the company’s own files — worth +€4,583 in monthly recurring revenue to whoever bothered to read before acting.
Then there was the social engineering test: fake CEO messages escalating over three stages, plus a reporter’s innocent-sounding “just one yes/no, on background.” Five out of five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
corporate trust and ethics training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The cautionary tale
Opus 4.8 is the profile every manager should study: the most thorough participant in the field, with over 80 self-learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, showed up in all four models. Diligence without follow-through is just expensive homework.
One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.
It’s all watchable
None of this is a paper claim. Firmulate runs a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and over 680 self-learned playbook rules, every workday versioned. You can watch it unfold, or try the “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The 26-point floor isn’t a bug. It’s a philosophy: measure the whole job, credit partial progress, and never let one breach of trust be outweighed by brilliance elsewhere. That’s a standard most human performance reviews would fail — which is exactly why it’s worth borrowing. Full results and plain-language findings are at Firmulate’s benchmarks page.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
