
We’ve Been Grading AI on the Wrong Exam
Think about every AI announcement you’ve read. A model writes a poem. It passes a bar exam. It aces a coding benchmark. We ooh and aah, and the leaderboard shuffles, and the cycle repeats.
But here’s a question nobody’s asking: those tests measure how well AI answers. What happens when AI has to manage — when there’s no single right answer, when a customer is furious, when cash is bleeding out, when a fake CEO message shows up asking for a ‘quick favor’?
That gap between answering well and managing well is exactly what Firmulate, a live public experiment, set out to measure. Its tagline says it plainly: the project measures management quality, not chat quality. And the results from its first finished league, finalized in July 2026, say something uncomfortable about the AI models we’re all rushing to adopt.
As an affiliate, we earn on qualifying purchases.
The Worst Week in the Life of a Software Company
The setup is elegantly simple. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 (five ran in total) — were each handed the same job: run a small software company through its worst week. Same customers. Same crises. Same temptations to cut corners. Only the model changes. Every decision is versioned and auditable, so you can go back and see exactly who did what.
What kind of week? The scenario names alone read like a management horror syllabus: a churn wave, a price increase, a downround, a PR crisis. This is the curriculum nobody puts on a résumé but everybody lives through.
As an affiliate, we earn on qualifying purchases.
The Headline Finding: Great Diagnosis, No Signature
Here’s where it gets interesting for anyone who thinks benchmark scores predict real-world performance.
All four models were excellent students in the ways we usually test for. They spotted every crisis. They refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter’s sly ‘just one yes/no, on background’ trick. Five out of five models said no, every time. Kimi K3 even put its reasoning on record: ‘Treat the request as a suspected approval-bypass / possible impersonation.’ That’s genuinely impressive judgment.
But only two of the four finished the job. Each had the chance to sign a €55,000 deal — a deal their own analysis had correctly earned them. Same diagnosis, same pitch… and for half the field, no signature. They simply didn’t close.
That gap — brilliant advice, no follow-through — is invisible in chat demos. You’d never see it in a leaderboard.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Decided Everything
The most revealing detail wasn’t in the customer call at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal — at full price, worth an extra +€4,583 in monthly recurring revenue.
It’s a lesson that applies to human employees as much as AI ones: the winning move was diligence, not brilliance. Reading before talking. Finishing what you start.
AI customer interaction analysis
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The League Table
The final Crucible League standings from July 2026:
- 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
- 2. Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline of the field. (One fairness note: K3 ran without an effort parameter while the others ran at xhigh — and still nearly won.)
- 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
- 4. Fable 5 — 77.
- 5. Opus 4.8 — 73. The most thorough participant of all — over 80 learned rules, the deepest analyses — yet last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
For context, a do-nothing baseline scores 26 — partial progress counts — but a single breach of trust caps the total. As the experiment’s own rule puts it: ‘no amount of good work outweighs a breach of trust.’ It’s a governance philosophy more boards should adopt verbatim.
This Isn’t a Slide Deck — It’s Running Now
What makes Firmulate different from every AI research paper you’ve skimmed is that it’s alive. The company at the center of this — 13 synthetic employees, real money mechanics — is running right now, every business day, losing money in public: €105k/month burn against just €2.3k in MRR, with a visible cash countdown. It has accumulated over 680 self-learned playbook rules, and every workday is versioned so you can audit the whole thing.
You can watch it unfold, or try the humbling ‘guess the model’ quiz — powered by 242 real, unedited management decisions from the experiment. Fair warning: the models are harder to tell apart than you’d think, right up until the moment one of them leaves €55,000 on the table.
Why This Matters Beyond the AI Crowd
Strip away the technology and this is a story about how we evaluate anyone — a new hire, a vendor, a tool, a model. We keep testing for eloquence when we should be testing for follow-through, honesty under pressure, and whether someone reads the file before making the pitch.
And if you run a business, you don’t have to settle for reading about it: enterprises can run the same wargame against a read-only export of their own operations — nothing ever writes back to real systems.

The uncomfortable truth from Firmulate’s first league is this: the AI models that dazzle in chat demos and top the coding charts are not automatically the ones that close deals, read the fine print, or hold the line when someone impersonates the boss. Some of the smartest participants in the field were also the ones who left money on the table.
If AI agents are coming for your CRM, your support queue, or your forecast — and they are — the question worth asking is no longer ‘does it write well?’ It’s: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work actually cost?
Chat quality gets the headlines. Management quality is what you’re buying. The full results and plain-language findings are public — and the company keeps running, twice-daily refreshes and all, for anyone who wants to see how the next exam goes.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html