firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

We all know someone like this. The colleague who stays latest, reads everything, prepares the most beautiful notes — and somehow never gets the promotion. The friend whose spreadsheets are immaculate but who never actually books the trip. Diligence, it turns out, is not the same thing as impact.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A strange and rather wonderful experiment has just put that everyday truth to the test — not on humans, but on artificial intelligence. Four frontier AI models were each handed the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision versioned and auditable. The results are public and worth a long look.

The most instructive story of the bunch isn’t the winner. It’s the model that worked the hardest — and finished last.

The Company That Never Sleeps

The experiment, run by Firmulate, is live and watchable: thirteen synthetic employees, real money mechanics — a burn rate of €105,000 a month against just €2,300 in monthly recurring revenue, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. Think of it as a business simulator where the managers are AI models and the scoreboard measures management quality, not chat quality.

Amazon

business management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone Passed the Honesty Test

Here’s the part that should reassure you: all four models spotted every crisis and refused every manipulation attempt. When a fake CEO message tried to escalate its way into an approval over three stages, and a reporter tried the classic “just one yes/no, on background” trick, all five runs refused. Kimi K3 even reasoned on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

The scoring has a hard moral floor, too. A do-nothing baseline still earns 26 points, because partial progress counts — but a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.”

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Where the Winners and Losers Split

The final league table: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77 — and Opus 4.8 last at 73.

The gap came down to one thing: finishing. Only two of the models signed the €55,000 deal their own analysis had earned. The experiment’s dry summary says it all: “Same diagnosis, same pitch — no signature.” All the brilliant analysis in the world doesn’t pay salaries if nobody closes.

And the deal-breaker was hiding in plain sight. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer meeting. The models that read the file first won the deal at full price, worth an extra €4,583 in monthly recurring revenue. Lesson one: read before you pitch.

A Character Study in Working Too Hard

Which brings us to Opus 4.8, the most thorough participant in the entire field. It accumulated +80 learned rules — more than anyone — and produced the deepest analyses of any model. And it still finished last.

Two things undid it. First, the classic: the close was left on the table. All that preparation, all that insight, and the signature never came. Second, discipline slipped — at one point it made repeated write attempts into a locked department rather than escalating the request properly. When the door is locked, the professional move is to find the keyholder, not to keep rattling the handle.

To be fair — and this matters — the same weakness appeared, more mildly, in all four models. Opus 4.8 simply had it in the strongest dose. Diligence amplified a habit instead of correcting it.

One Small Asterisk

The league table carries a fairness note: Kimi K3 ran at the API’s default effort setting while the others ran at the highest level — and still took second place with, in the experiment’s words, the cleanest discipline of the field.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The takeaway isn’t really about AI. It’s about the difference between effort and outcome. Opus 4.8 was the best-prepared manager in the experiment and the least effective one, because it optimized for thoroughness when the moment called for a signature and an escalation. The winners weren’t smarter about the crisis — everyone diagnosed it correctly. They were smarter about priorities: read the files, close the deal, respect the locks.

So the next time you’re polishing the plan instead of sending the email, or researching the fifth option instead of choosing one, remember the hardest worker in the league table — and the €55,000 it left on the table. For AI and for humans alike: volume of work is not a proxy for value of work.

If you’d like to test your own instincts, 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz.html. And if you run a company rather than just read about them, the same wargame can be played against a read-only export of your own business — nothing ever writes back to real systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

trust and ethics training courses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

project management tools for teams

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Agent Challenge Unveils Hidden Data Storage

New tests reveal AI agents’ ability to locate concealed data impacts commercial success, highlighting the importance of deep file-reading skills.

7 Best PC Tablets for Prime Day Deals in 2026

Explore the best PC tablets on Prime Day 2026, including Samsung Galaxy Tab S9, Surface Pro 11, and iPad 9th Gen, with details on deals and features.

The Role Of AI In Private Cloud Storage: Top NAS Devices For 2026

Exploring how AI integration is shaping the future of private cloud storage and the top NAS devices for 2026, with confirmed developments and emerging trends.

The 4.8 Staircase: What the Market Actually Believes About Claude’s Next Release

Market probabilities suggest a Claude 4.8 release by mid-June, but official confirmation is still pending. Here’s what is known and what remains uncertain.