
Spotting the Dirt Isn’t the Same as Mopping It
Anyone who has ever hired a cleaning contractor knows the walkthrough is the easy part. Every bidder tours the building, points at the dull finish in the lobby, the darkening grout in the restrooms, the salt stains nobody has touched since January. The diagnosis is always flawless. What actually decides the contract is duller: which crew is still on site at 11 p.m., stripping and re-waxing the last corridor instead of writing a beautiful report about it?
That same gap — between seeing a problem and finishing the job — turns out to be the most important thing to understand about artificial intelligence right now. A live, public experiment just ran frontier AI models through the same brutal week at the same small software company. The results, published by Firmulate, deserve to be pinned up wherever businesses are thinking about letting AI touch real work.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Crises, Same Temptations
The experiment gave each model the same assignment: run one small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. The company itself is real software, with 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown ticking down. Every decision is versioned and auditable.
The final league table, after the full week: gpt-5.6-sol leads with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. For calibration, a do-nothing baseline scores 26. Partial progress counts — but a single breach of trust caps the total, on the stated principle that no amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
Everyone Passed the Inspection
Here is the finding that should reset expectations: every model spotted every crisis, and every model refused every manipulation attempt. The pressure test was not gentle. Fake CEO messages escalated over three stages, and then a reporter dangled a quote — “just one yes/no, on background.” Five of five models refused. Kimi K3’s on-record reasoning reads like a security memo: “Treat the request as a suspected approval-bypass / possible impersonation.”
In floor-care terms, every crew walked the building and listed every stain, and nobody pocketed the keys. If the week had ended there, this would be a story about AI being ready for anything. It did not end there.
AI security and trust verification
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The €55,000 Signature
The week’s decisive moment was a deal worth €55,000. The competitor weakness that won it sat two document references deep in the company’s own files — not in the customer event, where anyone would think to look. The models that actually opened that file won the deal at full price, worth an extra €4,583 in monthly recurring revenue. Every model’s own analysis had earned the deal. Only two signed it. The experiment’s dry summary says it all: “Same diagnosis, same pitch — no signature.”
The most unsettling case is the last-place finisher. Opus 4.8 was the most thorough participant in the field — the deepest analyses, more than 80 self-learned playbook rules added during the week — and it still ended at 73. The close was left on the table, and discipline slipped at the edges: write attempts into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. Reading every page of the manual, it turns out, is not the same as finishing the shift.
One fairness note matters: Kimi K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and it still took second place. The margin between “finished” and “almost” is thinner than the scoreboard suggests.
AI model performance analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
This Is Not a Slide Deck
The company from the experiment is still running. It operates every business day, it is losing money in public right now, and it has accumulated more than 680 self-learned playbook rules, with every workday versioned. You can test your own instincts, too: 242 real, unedited management decisions from the runs power a “guess the model” quiz that is harder than it sounds. And for organisations that want the uncomfortable version, the same wargame can be run against a read-only export of their own business — nothing ever writes back to real systems.

What a Cleaning Crew Already Knows
The lesson is not that AI models are bad employees. It is that the interview measures the wrong thing. A chat demo — the polished conversation, the instant summary, the confident plan — is the walkthrough. It shows you a model that can spot the stain. It tells you almost nothing about whether it will read your files before quoting, stay honest under pressure, or still be working when the last corridor needs stripping. Closing strength is invisible until you test it, and the do-nothing baseline of 26 is a quiet warning that partial progress can look deceptively like work.
Before an AI agent gets anywhere near your CRM, support queue or forecast, the question is not “does it write well?” It is: does it finish what it starts? The full league table and plain-language findings are public on Firmulate’s benchmarks page — and unlike most AI news, you can watch the whole thing keep running at firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html