
We all know someone like this. The colleague who stays latest, reads everything, prepares the most beautiful notes — and somehow never gets the promotion. The friend whose spreadsheets are immaculate but who never actually books the trip. Diligence, it turns out, is not the same thing as impact.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
A strange and rather wonderful experiment has just put that everyday truth to the test — not on humans, but on artificial intelligence. Four frontier AI models were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision versioned and auditable. The results are public and worth a long look.
The most instructive story of the bunch isn’t the winner. It’s the model that worked the hardest — and finished last.
The Company That Never Sleeps
The experiment, run by Firmulate, is live and watchable: thirteen synthetic employees, real money mechanics — a burn rate of €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. Think of it as a business simulator where the managers are AI models and the scoreboard measures management quality, not chat quality.
business management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Everyone Passed the Honesty Test
Here’s the part that should reassure you: all four models spotted every crisis and refused every manipulation attempt. When a fake CEO message tried to escalate its way into an approval over three stages, and a reporter tried the classic “just one yes/no, on background” trick, all five runs refused. Kimi K3 even reasoned on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
The scoring has a hard moral floor, too. A do-nothing baseline still earns 26 points, because partial progress counts — but a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
Where the Winners and Losers Split
The final league table: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77 — and Opus 4.8 last at 73.
The gap came down to one thing: finishing. Only two of the models signed the €55,000 deal their own analysis had earned. The experiment’s dry summary says it all: “Same diagnosis, same pitch — no signature.” All the brilliant analysis in the world doesn’t pay salaries if nobody closes.
And the deal-breaker was hiding in plain sight. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer meeting. The models that read the file first won the deal at full price, worth an extra €4,583 in monthly recurring revenue. Lesson one: read before you pitch.
A Character Study in Working Too Hard
Which brings us to Opus 4.8, the most thorough participant in the entire field. It accumulated +80 learned rules — more than anyone — and produced the deepest analyses of any model. And it still finished last.
Two things undid it. First, the classic: the close was left on the table. All that preparation, all that insight, and the signature never came. Second, discipline slipped — at one point it made repeated write attempts into a locked department rather than escalating the request properly. When the door is locked, the professional move is to find the keyholder, not to keep rattling the handle.
To be fair — and this matters — the same weakness appeared, more mildly, in all four models. Opus 4.8 simply had it in the strongest dose. Diligence amplified a habit instead of correcting it.
One Small Asterisk
The league table carries a fairness note: Kimi K3 ran at the API’s default effort setting while the others ran at the highest level — and still took second place with, in the experiment’s words, the cleanest discipline of the field.

The takeaway isn’t really about AI. It’s about the difference between effort and outcome. Opus 4.8 was the best-prepared manager in the experiment and the least effective one, because it optimized for thoroughness when the moment called for a signature and an escalation. The winners weren’t smarter about the crisis — everyone diagnosed it correctly. They were smarter about priorities: read the files, close the deal, respect the locks.
So the next time you’re polishing the plan instead of sending the email, or researching the fifth option instead of choosing one, remember the hardest worker in the league table — and the €55,000 it left on the table. For AI and for humans alike: volume of work is not a proxy for value of work.
If you’d like to test your own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And if you run a company rather than just read about them, the same wargame can be played against a read-only export of your own business — nothing ever writes back to real systems.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
trust and ethics training courses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
project management tools for teams
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.