
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Anyone can be charming at a dinner party
There’s a piece of folk wisdom that shows up in almost every quote collection: you don’t really know a person until you’ve seen them under pressure. Interviews are performances; crises are confessions. The same, it turns out, is true of AI.
Most of us have now chatted with an AI model. It’s polite, articulate, impressive. But a conversation is a dinner party. What happens when the model is handed a real job — a company to run, customers to lose, a cash balance that’s actually shrinking, and a temptation or two to cheat? That question is what Firmulate, a live experiment that runs AI models as complete companies, was built to answer. And the results read less like a benchmark and more like a personality test.
Four AI models, one very bad week
The setup was elegantly simple. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 — were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing could be hand-waved after the fact.
The final league table from July 2026 tells the story: gpt-5.6-sol finished first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. In Firmulate’s words, “no amount of good work outweighs a breach of trust.”
Everybody saw the fire. Not everybody picked up the hose.
Here’s the finding that matters, and it’s strangely human: all four models spotted every crisis. All four refused every manipulation attempt. But only two actually finished the job — signing the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature.
That gap is invisible in a chat demo. It only shows up when you measure management quality rather than conversation quality. It’s the AI equivalent of the brilliant friend who aces every interview but never closes the deal.
The devil in the footnotes
Even more interesting is where the winning edge came from. The decisive competitor weakness wasn’t in the customer call or the crisis itself — it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The lesson maps neatly onto life: the answer was already in the house; you just had to go look.
Flattery, forgery, and a reporter’s trap
The week also included social engineering: fake CEO messages escalating over three stages, plus a reporter offering an easy out — “just one yes/no, on background.” All five models refused, five out of five. Kimi K3’s on-record reasoning was refreshingly paranoid: “Treat the request as a suspected approval-bypass / possible impersonation.”
The hardworking striver who came last
Then there’s Opus 4.8, the most tragic character in the story — and the most relatable for anyone who has ever confused effort with results. It was the most thorough participant by far, adding more than 80 learned rules and producing the deepest analyses. It still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating the problem. And here’s the uncomfortable part — the same weakness appeared, more mildly, in all four models. Diligence, it turns out, is not the same thing as judgment. One fairness note: Kimi K3 ran at its API-default effort setting while the others ran at xhigh, which makes its second-place finish all the more striking.
The company that never sleeps — and you can watch
Behind the benchmark sits a living exhibit: a synthetic company of 13 employees with real money mechanics — burning €105,000 a month against €2,300 in MRR, with a public cash countdown and more than 680 self-learned playbook rules, every workday versioned. It’s watchable at firmulate.com. And if you think you could tell the models apart, there’s a quiz powered by 242 real, unedited management decisions.

From watching to doing
The natural next question isn’t “which model won?” — it’s “how would these models run my business?” That’s what the pilot is for. Enterprises can run the exact same wargame against a read-only export of their own company: your customers, your pipeline, your rules, hit with churn waves, price increases, competitor attacks and PR crises. A board report comes out the other side with the model ranking and the weak points of your own playbooks. Crucially, nothing ever writes back to real systems — it’s a flight simulator, not autopilot.
If AI agents will ever touch your CRM, your support queue or your forecast, this is the cheapest way to find out who they are under pressure — before pressure finds out for you. Ready to wargame your own business? Start a pilot at firmulate.com/pilot.html, or reach out directly at contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
