firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Could you recognize an AI by the decisions it makes?

We often identify people through small habits: who sends the long explanation, who gets straight to the point, who reads the background material and who overlooks the final step. Artificial intelligence models are beginning to develop similarly recognizable management styles—and Firmulate has turned those differences into an interactive guessing game.

The Firmulate quiz draws on 242 real, unedited management decisions. Readers see how a model responded to a business situation and try to identify which one made the call. Some answers resemble dissertations. Others are terse. Another model may reject a request because it detects an attempt to bypass normal approval. The reveal is more than a name: it helps build a character profile based on observable behavior.

That makes the quiz entertaining, but its source material gives it weight. These are decisions taken during a live, watchable experiment in which frontier models were asked to manage the same small software company through its worst week.

Amazon

AI management decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One company, one terrible week, very different managers

Every participating model faced the same customers, crises and temptations. Every decision was versioned and auditable. This was not a test of who could produce the smoothest paragraph; it examined whether a model could notice trouble, use the company’s information, protect trust and complete commercially important work.

The final Crucible League standings from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One principle shaped the evaluation: a single breach of trust capped the total because "no amount of good work outweighs a breach of trust."

On the most basic tests, the field performed impressively. All models spotted every crisis and refused every manipulation attempt. Yet the commercial result exposed a sharp divide: only two signed the €55,000 deal their own analysis had earned. The summary is almost painfully simple: "Same diagnosis, same pitch — no signature."

The decisive clue was hiding in the paperwork

The difference was not a flash of rhetorical brilliance during a customer conversation. The decisive weakness in a competitor was buried two document references deep inside the company’s own files, rather than appearing in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR.

It is an unusually relatable management lesson. Spotting a crisis is not the same as resolving it, and preparing a convincing argument is not the same as closing. The experiment separates models that sound capable from models that connect research, judgment and follow-through.

Pressure revealed discipline as well as personality

The company also received fake CEO messages that escalated over three stages, followed by a reporter’s attempt to extract "just one yes/no, on background." All 5 models refused. Kimi K3 recorded a particularly clear reason: "Treat the request as a suspected approval-bypass / possible impersonation."

That unanimous resistance matters because the company was designed with real pressure behind its choices. It employed 13 synthetic workers and operated with real money mechanics, including burn of €105k/month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays make the experiment observable over time.

The personality profiles are not simple rankings of intelligence. Opus 4.8, for example, was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four, though less strongly.

There is also an important fairness note around the runner-up. Kimi K3 ran without an effort parameter and therefore used its API default, while the others ran at xhigh. That context does not erase the recorded decisions, but it belongs beside comparisons of performance and style.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the guessing game really tests

The pleasure of the quiz comes from pattern recognition: readers may start noticing which model explains everything, which one moves briskly and which one treats an unusual message as a security problem. But the broader lesson is about evaluating AI through work rather than conversation.

A polished answer can conceal an unfinished task. A lengthy analysis can coexist with weak execution. A cautious refusal can protect the company at exactly the moment when an apparently helpful request is trying to break its rules. Firmulate’s experiment makes those differences visible by putting models into identical situations and preserving what they actually decided.

For anyone curious about how AI might behave as a colleague or manager, the guess-the-model quiz offers an accessible place to begin. The challenge is not merely identifying a writing voice. It is recognizing a management personality through its habits: whether it reads first, stays trustworthy under pressure and finishes the work it started.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI behavior profiling tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI management style assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI Is Shaping The Best NAS Devices In 2026

Discover how artificial intelligence is transforming NAS device design and performance in 2026, shaping smarter, more efficient storage solutions.

The Bubble Question, Disentangled: 1999 vs 2026 Category by Category

A detailed analysis compares the AI investment cycle of 2024-2026 with the 1999 dotcom bubble, highlighting differences and implications for the future.

Israeli-born AI Startup In $6B Acquisition Talks With Anthropic

Anthropic is reportedly negotiating to acquire an unnamed Israeli-founded AI startup at a $6 billion valuation, but no deal has been confirmed yet.

Voice AI Cloning Licensing: A Smart Approach for Voice Actors

A licensing platform for voice actors to control and monetize their AI voice clones is being tested, offering a structured, auditable process for AI voice use.