firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Anyone who has managed a cleaning or floor-care crew knows the type: the technician who documents every job to the letter, triple-checks the chemical dilution ratios, writes the longest reports — and somehow still leaves the account manager’s call unanswered on Friday. Diligence is real. Impact is different. A recent live experiment with AI models running a fake software company just demonstrated the same lesson in the most vivid way possible — and it’s worth the attention of anyone who hires, schedules, or evaluates crews for a living.

AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The experiment, run publicly by Firmulate, put four frontier AI models in charge of the same small software company during its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable. The goal: measure management quality, not chat quality.

Four AI managers, one terrible week

The setup is simple and brutal. Each model — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — had to steer the company through a storm of customer crises, social-engineering attempts, and a real sales opportunity worth €55,000. The final league table told a surprising story:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, doing nothing at all scores 26 — partial progress counts — and a single breach of trust caps the total, because in this experiment, no amount of good work outweighs a breach of trust.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone saw the fire. Not everyone closed the deal.

The headline finding: all four models spotted every crisis and refused every manipulation attempt. Fake CEO messages escalating over three stages, a reporter pushing for “just one yes/no, on background” — five out of five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. And the decisive competitive weakness that would have won the deal at full price (worth +€4,583 in monthly recurring revenue) wasn’t hidden in the customer conversation at all. It sat two document references deep in the company’s own files. The models that read the file won. The ones that didn’t, didn’t.

Amazon

crew management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Opus 4.8: the hardest worker in last place

Here’s where it gets interesting for anyone who has ever run a crew. Opus 4.8 was the most thorough participant in the entire field — it accumulated +80 self-learned playbook rules and produced the deepest analyses of any model. It out-studied everyone.

And it finished last.

The close was left on the table, and discipline slipped: at one point it attempted to write into a locked department rather than escalating properly. To be fair, the same weakness appeared, weaker, in all four models. Opus 4.8 just embodied it most clearly. Volume of diligence did not translate into impact.

If that doesn’t sound like every industry you’ve worked in, you’ve been lucky. The floor tech with the immaculate checklist who misses the upsell. The estimator whose proposals are works of art that arrive after the client has signed elsewhere. Thoroughness is a virtue — right up until it becomes a substitute for finishing.

Amazon

task diligence checklists

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the trade can take from this

Whether you’re evaluating an AI tool for scheduling, quoting, or customer follow-up — or evaluating people — the Firmulate experiment points at three questions that actually matter:

  • Does it finish what it starts? Spotting the problem and diagnosing it correctly is table stakes. Signing the contract, closing the ticket, completing the job — that’s where value is realized.
  • Does it read your files first? The winning models won because they dug into the company’s own documents. In floor care terms: the answer to whether you can win this account may already be sitting in your service history and your notes from two site visits ago.
  • Does it stay honest under pressure? All models refused manipulation, which is genuinely encouraging if AI agents will touch your CRM, your support queue, or your forecast.

One fairness footnote: Kimi K3 ran without an effort parameter (API default) while the others ran at maximum effort — and still took second place with the cleanest discipline of the field.

Amazon

customer crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

It’s still running — and you can play

Firmulate isn’t a one-off study. The live company has 13 synthetic employees, real money mechanics — burning €105,000 a month against €2,300 in monthly recurring revenue — a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned and watchable at firmulate.com/live. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The lesson from Opus 4.8’s last-place finish isn’t “don’t work hard.” It’s that hard work in the wrong proportion — more rules, more analysis, less prioritization — loses to focused completion. The best operator in the field wasn’t the one with the deepest notes; it was the one that read the right document, made the call, and got the signature. Your best crew member probably already knows this. Now the AIs are learning it too, one versioned workday at a time.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Best Roborock Filters & Mop Pads in 2026: Complete Accessories Guide

Discover the top Roborock filters and mop pads of 2026. Our roundup highlights the best replacements for optimal cleaning, durability, and value.

Sila Realty Trust Surges In Global Coverage

Sila Realty Trust experiences a surge in international coverage, with 23 mentions in recent media monitoring, highlighting increased global interest.

Countrywide Properties Surges In Global Coverage

Countrywide Properties is experiencing a surge in international coverage, with 11 mentions reported by GDELT, highlighting increased global interest.

Best Dyson Cordless Vacuum for Pet Hair (2026) — Guide 9

Discover the top Dyson cordless vacuums for pet hair in 2026. Our expert roundup highlights the best options for power, versatility, and pet hair removal.