firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

We’ve Been Grading AI on the Wrong Exam

Think about every AI announcement you’ve read. A model writes a poem. It passes a bar exam. It aces a coding benchmark. We ooh and aah, and the leaderboard shuffles, and the cycle repeats.

But here’s a question nobody’s asking: those tests measure how well AI answers. What happens when AI has to manage — when there’s no single right answer, when a customer is furious, when cash is bleeding out, when a fake CEO message shows up asking for a ‘quick favor’?

That gap between answering well and managing well is exactly what Firmulate, a live public experiment, set out to measure. Its tagline says it plainly: the project measures management quality, not chat quality. And the results from its first finished league, finalized in July 2026, say something uncomfortable about the AI models we’re all rushing to adopt.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Worst Week in the Life of a Software Company

The setup is elegantly simple. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 (five ran in total) — were each handed the same job: run a small software company through its worst week. Same customers. Same crises. Same temptations to cut corners. Only the model changes. Every decision is versioned and auditable, so you can go back and see exactly who did what.

What kind of week? The scenario names alone read like a management horror syllabus: a churn wave, a price increase, a downround, a PR crisis. This is the curriculum nobody puts on a résumé but everybody lives through.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Headline Finding: Great Diagnosis, No Signature

Here’s where it gets interesting for anyone who thinks benchmark scores predict real-world performance.

All four models were excellent students in the ways we usually test for. They spotted every crisis. They refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter’s sly ‘just one yes/no, on background’ trick. Five out of five models said no, every time. Kimi K3 even put its reasoning on record: ‘Treat the request as a suspected approval-bypass / possible impersonation.’ That’s genuinely impressive judgment.

But only two of the four finished the job. Each had the chance to sign a €55,000 deal — a deal their own analysis had correctly earned them. Same diagnosis, same pitch… and for half the field, no signature. They simply didn’t close.

That gap — brilliant advice, no follow-through — is invisible in chat demos. You’d never see it in a leaderboard.

Amazon

AI business simulation models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact That Decided Everything

The most revealing detail wasn’t in the customer call at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal — at full price, worth an extra +€4,583 in monthly recurring revenue.

It’s a lesson that applies to human employees as much as AI ones: the winning move was diligence, not brilliance. Reading before talking. Finishing what you start.

Amazon

AI customer interaction analysis

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The League Table

The final Crucible League standings from July 2026:

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal. The complete performance.
  • 2. Kimi K3 — 93. The newcomer from Moonshot closed the deal too, with the cleanest discipline of the field. (One fairness note: K3 ran without an effort parameter while the others ran at xhigh — and still nearly won.)
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77.
  • 5. Opus 4.8 — 73. The most thorough participant of all — over 80 learned rules, the deepest analyses — yet last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

For context, a do-nothing baseline scores 26 — partial progress counts — but a single breach of trust caps the total. As the experiment’s own rule puts it: ‘no amount of good work outweighs a breach of trust.’ It’s a governance philosophy more boards should adopt verbatim.

This Isn’t a Slide Deck — It’s Running Now

What makes Firmulate different from every AI research paper you’ve skimmed is that it’s alive. The company at the center of this — 13 synthetic employees, real money mechanics — is running right now, every business day, losing money in public: €105k/month burn against just €2.3k in MRR, with a visible cash countdown. It has accumulated over 680 self-learned playbook rules, and every workday is versioned so you can audit the whole thing.

You can watch it unfold, or try the humbling ‘guess the model’ quiz — powered by 242 real, unedited management decisions from the experiment. Fair warning: the models are harder to tell apart than you’d think, right up until the moment one of them leaves €55,000 on the table.

Why This Matters Beyond the AI Crowd

Strip away the technology and this is a story about how we evaluate anyone — a new hire, a vendor, a tool, a model. We keep testing for eloquence when we should be testing for follow-through, honesty under pressure, and whether someone reads the file before making the pitch.

And if you run a business, you don’t have to settle for reading about it: enterprises can run the same wargame against a read-only export of their own operations — nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The uncomfortable truth from Firmulate’s first league is this: the AI models that dazzle in chat demos and top the coding charts are not automatically the ones that close deals, read the fine print, or hold the line when someone impersonates the boss. Some of the smartest participants in the field were also the ones who left money on the table.

If AI agents are coming for your CRM, your support queue, or your forecast — and they are — the question worth asking is no longer ‘does it write well?’ It’s: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work actually cost?

Chat quality gets the headlines. Management quality is what you’re buying. The full results and plain-language findings are public — and the company keeps running, twice-daily refreshes and all, for anyone who wants to see how the next exam goes.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Waves, Not a Wall: Inside DeepMind’s Map From AGI to Superintelligence

DeepMind researchers publish a detailed framework outlining pathways from human-level AI to superintelligence, highlighting scaling, paradigm shifts, and challenges.

Europe Regulated the Interface and Forgot to Build the Engine

Europe has regulated the AI interface but failed to develop the underlying technology, risking its position in the global AI race amid rising competition.

7 Best PC Routers for Prime Day Deals in 2026

Explore the best PC router deals for Prime Day 2026, including WiFi 7 models, wired options, and gaming routers, tailored for different user needs.

Unveiling The Best AI 4K Monitors For Work And Gaming In 2026

Discover the top AI-enhanced 4K monitors for work and gaming in 2026, featuring the latest models, specs, and buying tips for every user type.