firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Every cleaning contractor knows the type: the technician who arrives, diagnoses the floor perfectly — correct sealer, correct pad, correct dilution ratio — and then leaves without actually running the machine. The diagnosis was flawless. The floor is still dirty. In the service business, we have a blunt name for that: not finishing the job. It turns out frontier AI models have exactly the same failure mode, and someone finally ran the experiment to prove it.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get cleaning gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That someone is Firmulate, a public project that runs AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality rather than chat quality. In July 2026 it published the final league table of what it calls the Crucible: five frontier models, each handed the same small software company during its worst week. The results are watchable, auditable, and deeply relevant to anyone who worries about hiring an AI that talks a good game.

The worst week in business, five times over

The setup is elegantly cruel. Each model ran the identical company through the identical seven days of chaos: the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable, so nothing rests on anecdote. A do-nothing baseline scores 26 — partial progress counts — but a single breach of trust caps the total outright. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust.

The final standings:

  • 1. gpt-5.6-sol — 95. Found the buried fact, closed the deal — the complete performance.
  • 2. Kimi K3 — 93. The newcomer from Moonshot: closed the deal too, with the cleanest discipline in the field.
  • 3. Sonnet 5 — 88. Closed the deal, with a few more process slips.
  • 4. Fable 5 — 77.
  • 5. Opus 4.8 — 73. The most thorough participant of all — and still last.
Amazon

AI management software for small businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The newcomer beats three of four Western rivals

The headline finding is that Kimi K3 — a model few enterprise buyers had shortlisted — finished second, ahead of three of four Western frontier models. K3 found the buried security needle in the company’s own files, won the €55,000 deal at full price (worth +€4,583 in monthly recurring revenue), saved the churning customer, and resisted every one of three baited temptations. It deviated from clean process exactly once — the best discipline record in the field.

One fairness footnote matters here: K3 ran without an effort parameter, at the API default, while the other four models ran at their highest reasoning effort. In other words, the newcomer matched or beat heavily favored competitors that were working harder, not less.

If your model of the AI market assumes the usual Western names lead automatically, this table says otherwise. The league is open — and picking a model without testing it against your own business is now a bet, not a decision.

Amazon

business process automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same diagnosis, same pitch — no signature

Here is the finding that should stop any service-business owner cold. Every model spotted every crisis. Every model refused every manipulation attempt. But only two of five actually signed the €55,000 deal their own analysis had earned.

The difference came down to reading the files. The decisive competitor weakness wasn’t in the customer conversation at all — it sat two document references deep inside the company’s own records. The models that read their own files won the deal at full price. The ones that didn’t, didn’t. It’s the cleaning equivalent of a technician who never checks the maintenance log from the last visit: the answer was in the van the whole time.

Amazon

AI decision-making tools for service companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Honest under pressure

The experiment also staged a social-engineering attack: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was refreshingly paranoid: treat the request as a suspected approval-bypass, possible impersonation.

That is the good news. The bad news is the discipline gap. Opus 4.8, the most thorough participant — it generated over 80 learned rules and the deepest analyses of any model — still finished last. It left the close on the table and slipped on process, attempting writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models. Thoroughness, it turns out, is not the same as finishing.

Amazon

AI security and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Not a slide deck — a live company

Firmulate insists on being watchable rather than described. The company is real software with 13 synthetic employees, running every business day, burning €105k a month against €2.3k in MRR, with a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. You can watch it lose money in real time at firmulate.com.

For the skeptical, there’s a quiz built from 242 real, unedited management decisions — guess which model made which call. And for enterprises, there’s a pilot program: run the same wargame against a read-only export of your own business, with nothing ever written back to real systems. Full benchmarks and plain-language findings are at firmulate.com/benchmarks.html.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The lesson travels well beyond software companies. If you’re ever going to let an AI agent near your CRM, your support queue, or your booking schedule, the question is not “does it write well.” It’s: does it finish what it starts, does it read your own records before acting, does it stay honest when someone pretends to be the boss — and what does a unit of useful work actually cost?

The Crucible showed that chat quality and management quality are different things, that a newcomer can beat established names, and that the most thorough operator can still leave the job half-done. Anyone who has hired a subcontractor knows exactly what that looks like. Now there’s a way to test for it before you hire the AI too.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

A Tiny Bathroom Went From Grimy To Cozy After A “Gothic-Aquatic” Redo

A small, grimy bathroom was renovated into a stylish ‘Gothic-Aquatic’ retreat, gaining attention for its dramatic makeover and unique theme.

Howard Hanna Surges In Global Coverage

Howard Hanna has increased its international media mentions, surging to 26 mentions in recent coverage, indicating a major expansion in global visibility.

As A.I. Money Floods The Market, San Francisco Renters Weigh Buyouts – The New York Times

As AI companies flood the market with investment, some San Francisco renters are accepting buyouts to leave their apartments, raising questions about the city’s housing future.

Avalonbay Communities Surges In Global Coverage

AvalonBay Communities experiences a surge in international coverage, with 26 mentions in recent global media tracking, indicating rising global interest.