
Get the little things that make your day delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Hiring Used to Be About Résumés. Now It’s About Your Worst Week.
Most of life’s big decisions come down to the same question: how does this person behave when things go wrong? We accept the job, marry the person, trust the friend — not because of the polished pitch, but because of what we saw under pressure. For a growing number of teams, the next hire isn’t a person at all. It’s an AI model that will touch your customers, your inbox, your money. And pressure, it turns out, is exactly where the surprises live.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Crisis, Five Different Minds
That’s the idea behind Firmulate, a public experiment that runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. In its Crucible league, finished in July 2026, five frontier models were each handed the same small software company on its worst week: same customers, same emergencies, same chances to cut corners. Only the model changed. Every decision was versioned and auditable.
The final table: gpt-5.6-sol scored 95. Moonshot’s Kimi K3 scored 93. Sonnet 5 took 88, Fable 5 got 77, and Opus 4.8 landed last at 73. For context, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
The Story of the Week: K3, the Unknown Who Closed
K3’s run reads like a quiet underdog story. It found the buried security needle hidden in the company’s own files — not in the customer conversation, but two document references deep. It won the €55,000 deal at full price, worth +€4,583 in monthly recurring revenue. It saved the churning customer. And when fake CEO messages came escalating in three stages, followed by a reporter’s sly “just one yes/no, on background” trick, K3 refused everything — its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the whole week, it logged just one deviation: the cleanest discipline in the field.
The buried fact turned out to be the week’s hidden separator. The decisive competitor weakness sat in the company’s own documents, not the customer event. The models that actually read the file won the deal at full price.
AI security document analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Everyone Diagnosed It. Only Some Finished It.
Here’s the finding that should stop any buyer mid-purchase: all five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in a chat demo.
Then there’s Opus 4.8, the cautionary tale: the most thorough participant in the field, with 80 additional learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.
enterprise AI risk assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
It’s All Watchable — and Playable
The company at the heart of this is live software, not a slide deck. It employs 13 synthetic employees, burns €105k a month against €2.3k in MRR, has a public cash countdown, over 680 self-learned playbook rules, and rebuilds its public site twice a day — watchable at firmulate.com. You can also try beating the models yourself: 242 real, unedited management decisions power a guess-the-model quiz, and full plain-language findings sit on the benchmarks page. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
Fairness note: Kimi K3 ran without an effort parameter (API default) while the other models ran at xhigh.

The League Is Open
The comfortable assumption that the big Western labs sit safely on top took a hit this month. A newcomer from Moonshot beat three of four Western frontier models at the most human of tasks: running a company through its worst week. If you’re picking a model on brand reputation alone, you’re not making a decision — you’re placing a bet. Test it on your own worst week first. The tools to do exactly that now exist, and they’re open to anyone with a browser.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
