firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Anyone in the floor care business knows the difference between starting a job and finishing it. A crew can strip the wax, buff to a shine, and still fail the contract because nobody did the edges. The client sees the corners. That instinct — that a job half-done is a job failed — is exactly the logic behind a surprising new benchmark that grades AI models on how they run a company, not how they chat.

Before you orderOffer from Amazon

Get cleaning gear delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The experiment, run by a firm called Firmulate, handed four frontier AI models the same small software company and the same catastrophic week: same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable. The final league table has a twist that should feel familiar to anyone who has ever graded a cleaning contract: the top model, gpt-5.6-sol, scored 95 out of 100 — and the whole system is built to distrust a perfect 100.

Why a Do-Nothing Manager Gets 26, Not 0

The strangest number in the whole benchmark is the floor: a baseline run that does essentially nothing still scores 26 points. That is not a bug. It is a statement about how honest grading works.

In floor care terms: if a contractor shows up, surveys the site, puts up the wet-floor signs and identifies the right chemicals but never touches the buffer, they have produced partial value. A scoring system that gives them zero is lying about what happened. Firmulate’s approach counts partial progress — the diagnosis, the correct identification of the job — because in real operations, knowing what needs doing is genuinely worth something. The do-nothing baseline of 26 exists so every score above it has meaning: it is what a model earned beyond merely showing up.

Amazon

automatic floor buffer machine

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Breach of Trust Caps the Whole Grade

The second design choice is harsher. A single breach of trust — dishonesty, deception, cutting a corner that betrays the client — caps the model’s total score, no matter how brilliant the rest of the work. The published reasoning is blunt: “no amount of good work outweighs a breach of trust.”

This is the cleaning industry’s oldest rule, codified. A crew can leave the marble immaculate, but if they were caught using a client’s supply closet for personal gain, that contract is over. Firmulate grades AI the way a facility manager grades a vendor: competence matters, but trust is the ceiling on everything else.

Amazon

wet floor safety signs

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Results: Everyone Spotted the Mess, Only Two Finished the Job

What happened when the models actually ran the company? The headline finding is eerie in how it mirrors human crews:

  • All four models spotted every crisis.
  • All four refused every manipulation attempt, including a three-stage fake-CEO impersonation and a reporter’s “just one yes/no, on background” trick — 5 of 5 refusals across the field. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
  • Only two models signed the €55,000 deal their own analysis had earned. The benchmark’s verdict on the others: “Same diagnosis, same pitch — no signature.”

That is the mopped-floor-with-dirty-edges problem exactly. The models could diagnose, prepare and present — but when it came to closing, some left the job unfinished. And that gap is invisible in chat demos, which is precisely why Firmulate built a wargame instead.

The final standings from July 2026’s Crucible League: gpt-5.6-sol first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won.

Amazon

floor cleaning chemicals set

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact That Won the Deal

How did the winners close at full price, worth +€4,583 in monthly recurring revenue? Not through the customer conversation at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the paperwork found it and won. The models that skimmed did not.

Again, the floor care parallel writes itself: the notes from last year’s site inspection, buried in the file cabinet, tell you exactly why this client left the last vendor. Read the file, win the contract. Skip it, and you give the same pitch as everyone else.

Amazon

professional floor cleaning tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

When Thoroughness Isn’t Enough

The most instructive profile is Opus 4.8: the most thorough participant in the field, generating over 80 learned rules and the deepest analyses — and still last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four models. Effort is not the same as completion.

It’s Live, and You Can Test Yourself

The experiment is not a static paper. Firmulate runs a live, watchable company: 13 synthetic employees, real money mechanics — burning €105k a month against just €2.3k in monthly recurring revenue — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live.

There is also a quiz built from 242 real, unedited management decisions, where you guess which model made which call (firmulate.com/quiz.html). And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The lesson for anyone who manages crews, contracts or vendors: the future of judging AI will look a lot like judging a cleaning crew. Did it finish the whole floor, or just the visible middle? Did it read the file before pitching? Does one betrayal cap the entire grade, regardless of shine? Firmulate’s answer — with its honest 26-point floor, its trust cap, and its suspicion of a perfect 100 — is the closest thing yet to a facility-manager’s checklist for AI. And in the first running of it, the winner was not the flashiest talker. It was the model that read everything, closed the deal, and left nothing on the table. Full results are public here.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

I Have A Leg Injury, And This Extendable Spin Scrubber Has Been A Lifesaver

A person with a leg injury finds relief using an extendable spin scrubber, highlighting its potential as an assistive tool for mobility-impaired individuals.

West Elm Just Gave Its Holiday Collection A Very New York Twist

West Elm unveils a holiday collection inspired by New York City, blending urban style with festive design elements, available now.

The AI Detail That Closed a €55,000 Deal — And What Floor Care Pros Already Know About Reading the Fine Print

All four AI models diagnosed the problem; only two read the paperwork and closed the €55k deal. The prep-work lesson floor care pros already know.

6 Top-Rated Wet-Dry Vacuums That Will Clean Up After Pets And Kids Effortlessly

Discover the six best-rated wet-dry vacuums ideal for cleaning up after pets and children, offering effortless and efficient mess removal.