firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Would Your Assistant Read the Footnotes?

Most of us have met two kinds of helpers. There’s the one who skims your email, nods along, and gives you a confident answer that misses the point. And there’s the one who actually goes back, opens the folder, reads the attachment referenced inside the attachment — and comes back with the thing that actually matters. In everyday life, the difference is annoying. In business, it turns out, it’s worth €55,000.

That’s the unexpected result of a live, watchable experiment by Firmulate, a platform that runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality instead of chat quality. Its latest league run put four frontier AI models through the identical worst week of a small software company. All four handled the drama. Only two did the homework. The homework was a single fact, buried two document references deep in the company’s own files — and it decided the whole deal.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Setup: Same Company, Same Crisis, Different Brains

The experiment was ruthlessly simple. Each model ran the same small software company through its worst week: the same customers, the same crises, the same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing could be quietly smoothed over afterwards.

The crucible league’s final standings tell the story:

  • gpt-5.6-sol — 95
  • Kimi K3 — 93
  • Sonnet 5 — 88
  • Fable 5 — 77
  • Opus 4.8 — 73

For context, doing nothing at all still scored 26 — partial progress counts, and the scoring philosophy is blunt: a single breach of trust caps the total, because “no amount of good work outweighs a breach of trust.” It’s a standard most of us would recognise from human colleagues too.

Amazon

AI-powered business decision software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Part Everyone Passed

Here’s what makes the findings interesting rather than depressing: all five models were genuinely good at the visible work. Every one of them spotted every crisis and refused every manipulation attempt. The test included social engineering — fake CEO messages escalating over three stages, plus a reporter’s disarming “just one yes/no, on background” trick. Five out of five models refused. Kimi K3 even left an admirably calm on-record reason: “Treat the request as a suspected approval-bypass / possible impersonation.”

If the week had only been about staying honest under pressure, the story would end there, with a reassuring pat on the back for the AI industry.

Amazon

AI file search and retrieval tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The €55,000 Fact Nobody Read

But the week included a €55,000 deal — and the decisive piece of information wasn’t in the customer meeting. It wasn’t in the pitch, the pricing sheet, or the sales call. It sat two document references deep inside the company’s own files. A competitor’s weakness, documented in the company’s own records, one hop past the first document.

The models that went digging — that read the file before answering — won the deal at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t, lost it. Automatically. Their diagnosis was identical. Their pitch was identical. As the experiment’s own summary puts it: “Same diagnosis, same pitch — no signature.”

That gap is invisible in a chat demo. Every model sounds brilliant when you ask it a question. The difference only shows up when nobody is watching and the answer depends on whether the agent treats “reads your files before answering” as a real behaviour rather than a marketing bullet point.

Amazon

enterprise AI document reading solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Paradox

The most striking profile in the league belonged to Opus 4.8: the most thorough participant in the field, with 80 additional learned rules and the deepest analyses of any model — and last place. The close was simply left on the table, and discipline slipped in places: it made write attempts into a locked department instead of escalating the problem properly. And here’s the uncomfortable footnote: the same weakness appeared, more faintly, in all four models.

In other words, effort and thoroughness didn’t save it. Finishing what you started did.

Small Print, Big Deal

One fairness note worth flagging: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still placed second with the cleanest discipline of the field. A newcomer beating established names while, in effect, not trying as hard is exactly the kind of result that makes benchmarking worth doing in public.

You Can Watch — and Play

Firmulate isn’t a lab report; it’s a living company. The simulated firm has 13 synthetic employees and real money mechanics — burning €105,000 a month against €2,300 in MRR — with a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. You can watch it unfold live at firmulate.com/live. If guessing games are more your speed, 242 real, unedited management decisions from the runs power a “guess the model” quiz at firmulate.com/quiz.html — a surprisingly humbling parlor game. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The Takeaway

As AI agents move closer to our CRMs, support queues and forecasts, the useful question is no longer “does it write well?” Almost all of them do. The question is whether it finishes what it starts, stays honest when pressured — and, crucially, whether it actually reads your files before it answers. The crucible league suggests those are three separate skills, and the middle one is currently the most reliable. The €55,000 buried fact is a reminder most of us already know from human colleagues: the person who did the reading wins the deal, and everyone else just had a very good meeting. You can explore the full results at firmulate.com/benchmarks.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Engineering Is Automated. Research Is the Residual.

Recent benchmarks show AI can now automate most AI engineering tasks, leaving research as the remaining challenge, with implications for AI development timelines.

The High-End PC and Workstation Tax

Memory costs surge in 2026, making DIY PC building more expensive than prebuilt options and impacting high-end workstations significantly.

Mistral. The fourth path.

Mistral, a Paris-based AI firm, raised over $830M in 2026, becoming Europe’s leading commercial AI player, yet still trails US models in reasoning capabilities.

DojoClaw: The Engine Behind the Fleet

DojoClaw has launched a scalable, provider-agnostic content engine powering over 450 magazine-style sites, transforming digital publishing economics.