firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A different kind of workplace drama

Some people begin the week with a motivational quote. Firmulate begins with a software company trying to stay alive.

The public experiment has the ingredients of an unusually transparent business story: 13 synthetic employees, burn of €105,000 a month, just €2,300 in monthly recurring revenue, and a cash countdown anyone can follow. Its workforce has accumulated more than 680 self-learned playbook rules, while every workday is versioned for later inspection.

This is not a polished simulation presented after the fact. The company is real software, its money mechanics are real, and the struggle is continuing in public. Visitors can watch the company live, turning ordinary business activity into an unfolding account of decisions, setbacks and survival.

That makes Firmulate a particularly extreme example of building in public. Most companies reveal selected milestones. This one exposes the gap between recognizing a problem and actually doing what is necessary to solve it.

Amazon

business decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What happens when frontier models become management

Firmulate’s Crucible League put frontier AI models in charge of the same small software company during its worst week. Each participant encountered the same customers, crises and temptations. Every decision was versioned and auditable.

The final July 2026 standings placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. However, a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The encouraging result was that all models noticed every crisis and rejected every attempt at manipulation. The more revealing result was that only two signed the €55,000 deal their own work had earned. Firmulate summarizes that performance gap neatly: “Same diagnosis, same pitch — no signature.”

For a general audience, this may be the most relatable part of the experiment. Knowing what should happen is not the same as finishing the job. A manager can understand a customer, prepare the right argument and still fail at the moment when action matters.

The detail hidden in plain sight

The decisive clue was not contained in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that followed those references found the information and won the deal at full price, adding €4,583 in monthly recurring revenue.

That finding makes the experiment feel less like a test of clever conversation and more like a test of workplace habits. The winning behavior was simple but demanding: read the available material, follow the trail and use the evidence when it counts.

Pressure without surrendering trust

The models also faced fake CEO messages that escalated across three stages, followed by a reporter’s attempted shortcut: “just one yes/no, on background.” All 5 models refused.

Kimi K3’s recorded explanation was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response matters because social pressure in business rarely arrives with a convenient warning label. It can sound urgent, authoritative or harmless. Here, every participant maintained the boundary.

There is a fairness qualification in K3’s strong second-place result. K3 ran using its API default because it had no effort parameter, while the other participants ran at xhigh. The distinction does not erase the outcome, but it is important context when comparing the field.

When thoroughness becomes its own trap

Opus 4.8 presents the most interesting cautionary portrait. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating the problem.

A weaker form of that same discipline problem appeared in all four other models. The lesson is not that careful analysis lacks value. It is that analysis, procedural judgment and completion must work together. A long, thoughtful trail can still end short of the result.

Those interested in the human texture behind the experiment can also read what the synthetic employees say. Their words help turn an abstract benchmark into something closer to a workplace chronicle.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

workplace trust training courses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that makes unfinished work visible

Firmulate’s public experiment offers a useful counterweight to effortless-looking AI demonstrations. Its synthetic workforce can identify crises, resist manipulation and produce deep analysis. Yet the league also shows how much depends on mundane habits: reading the files, respecting boundaries, escalating correctly and completing a hard-won sale.

Meanwhile, the live company continues burning €105,000 a month against €2,300 in monthly recurring revenue. That imbalance gives every workday narrative pressure. The public countdown is not decorative; it is the clock behind every decision.

For readers accustomed to lifestyle advice and memorable quotations, the story carries a practical message: competence is not merely having the right answer. It is following through without sacrificing trust. Firmulate has made that lesson watchable, versioned and unusually difficult to hide.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

employee training playbook templates

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

business crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Slow To Welcome AI, Hard To See It Leave

Analysis of how enterprise incumbents are slow to adopt AI but remain resistant to displacement, shaping the future of enterprise AI dynamics.

7 Best PC Motherboards for Prime Day Deals in 2026

Discover the best PC motherboard deals for Prime Day 2026, including options for AM4 and AM5 platforms, with insights on features and upgrade paths.

Create Content Smarter With These 12 AI Tools In 2026

Discover the top 12 AI tools for content creation in 2026, designed to streamline research, drafting, publishing, and moderation for diverse content needs.

2026’S WiFi 7 Routers With The Best AI Features

Explore the leading WiFi 7 routers of 2026 featuring advanced AI capabilities, high speeds, and improved coverage for demanding households.