
Anyone who has managed a cleaning or floor-care crew knows the type: the technician who documents every job to the letter, triple-checks the chemical dilution ratios, writes the longest reports — and somehow still leaves the account manager’s call unanswered on Friday. Diligence is real. Impact is different. A recent live experiment with AI models running a fake software company just demonstrated the same lesson in the most vivid way possible — and it’s worth the attention of anyone who hires, schedules, or evaluates crews for a living.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The experiment, run publicly by Firmulate, put four frontier AI models in charge of the same small software company during its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable. The goal: measure management quality, not chat quality.
Four AI managers, one terrible week
The setup is simple and brutal. Each model — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — had to steer the company through a storm of customer crises, social-engineering attempts, and a real sales opportunity worth €55,000. The final league table told a surprising story:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
For context, doing nothing at all scores 26 — partial progress counts — and a single breach of trust caps the total, because in this experiment, no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
Everyone saw the fire. Not everyone closed the deal.
The headline finding: all four models spotted every crisis and refused every manipulation attempt. Fake CEO messages escalating over three stages, a reporter pushing for “just one yes/no, on background” — five out of five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. And the decisive competitive weakness that would have won the deal at full price (worth +€4,583 in monthly recurring revenue) wasn’t hidden in the customer conversation at all. It sat two document references deep in the company’s own files. The models that read the file won. The ones that didn’t, didn’t.
As an affiliate, we earn on qualifying purchases.
Opus 4.8: the hardest worker in last place
Here’s where it gets interesting for anyone who has ever run a crew. Opus 4.8 was the most thorough participant in the entire field — it accumulated +80 self-learned playbook rules and produced the deepest analyses of any model. It out-studied everyone.
And it finished last.
The close was left on the table, and discipline slipped: at one point it attempted to write into a locked department rather than escalating properly. To be fair, the same weakness appeared, weaker, in all four models. Opus 4.8 just embodied it most clearly. Volume of diligence did not translate into impact.
If that doesn’t sound like every industry you’ve worked in, you’ve been lucky. The floor tech with the immaculate checklist who misses the upsell. The estimator whose proposals are works of art that arrive after the client has signed elsewhere. Thoroughness is a virtue — right up until it becomes a substitute for finishing.
As an affiliate, we earn on qualifying purchases.
What the trade can take from this
Whether you’re evaluating an AI tool for scheduling, quoting, or customer follow-up — or evaluating people — the Firmulate experiment points at three questions that actually matter:
- Does it finish what it starts? Spotting the problem and diagnosing it correctly is table stakes. Signing the contract, closing the ticket, completing the job — that’s where value is realized.
- Does it read your files first? The winning models won because they dug into the company’s own documents. In floor care terms: the answer to whether you can win this account may already be sitting in your service history and your notes from two site visits ago.
- Does it stay honest under pressure? All models refused manipulation, which is genuinely encouraging if AI agents will touch your CRM, your support queue, or your forecast.
One fairness footnote: Kimi K3 ran without an effort parameter (API default) while the others ran at maximum effort — and still took second place with the cleanest discipline of the field.
customer crisis management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
It’s still running — and you can play
Firmulate isn’t a one-off study. The live company has 13 synthetic employees, real money mechanics — burning €105,000 a month against €2,300 in monthly recurring revenue — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned and watchable at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

The lesson from Opus 4.8’s last-place finish isn’t “don’t work hard.” It’s that hard work in the wrong proportion — more rules, more analysis, less prioritization — loses to focused completion. The best operator in the field wasn’t the one with the deepest notes; it was the one that read the right document, made the call, and got the signature. Your best crew member probably already knows this. Now the AIs are learning it too, one versioned workday at a time.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.