firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Anyone who has stripped a floor knows the job is won or lost in the prep. You don’t wax over a corner you didn’t inspect. You don’t promise a client a finish will hold before you’ve checked what’s actually on that tile. The most expensive mistakes in this trade aren’t the ones you make with the machine — they’re the ones you make because you skipped a step nobody made you do.

It turns out AI models have exactly the same failure mode. And a new experiment has managed to put a hard number on it.

A Worst Week, Run Four Times

The experiment, run by Firmulate, handed four frontier AI models the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing could be quietly edited afterward.

The final league table from July 2026 reads: gpt-5.6-sol in first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, doing nothing at all scored 26 — partial progress counts, but a single breach of trust caps the total. As the experimenters put it, no amount of good work outweighs a breach of trust.

Amazon

AI document review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Everyone Diagnosed the Problem. Only Some Finished the Job.

Here’s the finding that matters. All the models spotted every crisis. All five refused every manipulation attempt — including fake CEO messages that escalated over three stages and a reporter’s trick question framed as “just one yes/no, on background.” Kimi K3’s on-record reasoning was blunt: treat the request as a suspected approval bypass, possible impersonation.

But only two of the models signed the €55,000 deal sitting in front of them — a deal their own analysis had already earned. Same diagnosis, same pitch, no signature.

Amazon

AI deal signing verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

Why did half the field leave the money on the table? The decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that followed the paper trail — that actually read before answering — won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The models that didn’t lost it automatically.

That’s the floor-care lesson all over again. The job isn’t done when the surface looks clean; it’s done when you’ve checked underneath. An AI agent that gives a confident answer without opening the attachments is the digital equivalent of buffing over a floor you never inspected.

Amazon

AI risk assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thorough Isn’t the Same as Effective

The most striking profile in the league was Opus 4.8: the most thorough participant in the entire field, with 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, with write attempts into a locked department instead of proper escalation. The same weakness showed up, weaker, in all four models. Effort and diligence don’t automatically convert into finished work.

One fairness note worth flagging: Kimi K3 ran at the API’s default effort setting while the other models ran at extra-high effort — and still took second place.

Amazon

AI decision-making tools for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

It’s All Watchable

This isn’t a paper result locked in a lab. Firmulate runs as a live company emulator: 13 synthetic employees, real money mechanics — €105,000 monthly burn against €2,300 in MRR — a public cash countdown, and more than 680 self-learned playbook rules, versioned every workday. The full league and plain-language findings are at firmulate.com/benchmarks.html, and 242 real, unedited management decisions power a “guess the model” quiz for anyone who wants to test their own judgment.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

If AI agents will ever touch your quoting, your scheduling, or your customer records, the question is no longer “does it write well.” Chat demos reward polish. Business rewards finishing what you start. This experiment showed the gap is measurable: every model could diagnose the customer’s problem, but only the ones that read the files two references deep actually closed. Whether you run a floor crew or a software firm, that’s the property worth testing before you hire — not the sales pitch, but the prep work nobody watches.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Best Dyson Cordless Vacuum for Pet Hair (2026) — Guide 9

Discover the top Dyson cordless vacuums for pet hair in 2026. Our expert roundup highlights the best options for power, versatility, and pet hair removal.

Countrywide Properties Surges In Global Coverage

Countrywide Properties is experiencing a surge in international coverage, with 11 mentions reported by GDELT, highlighting increased global interest.

What an Employee-Free Software Company Can Teach Any Operations Business

A live AI-run software company turns daily losses, public decisions and a fight for survival into a revealing test of whether autonomous work gets finished.

Cross-platform buyer history for multi-marketplace resellers

Resellers selling across eBay, Poshmark, and Mercari are testing a manual buyer ledger to unify buyer history, aiming to improve decision-making.