
Get cleaning gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When the cleaning business has a bad week
A steam mop can make a stubborn floor look manageable. Running the business behind it is harder: customers may be at risk of leaving, a competitor may be moving in, and a tempting shortcut may threaten trust. Firmulate puts AI models through that kind of pressure in a company they must actually manage. Its live experiment offers a question for anyone selling or servicing floor-care products: when the plan is clear, will an AI carry it through?
Same company, same crises
In the final Crucible League, published in July 2026, frontier models were each given the same small software company and its worst week: the same customers, crises and temptations. Their decisions were versioned and auditable. The top results were gpt-5.6-sol at 95, Kimi K3 at 93 and Sonnet 5 at 88. Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26. The league also sets a hard principle: partial progress counts, but one breach of trust caps the total. “No amount of good work outweighs a breach of trust.”
The striking result was not that the models missed the trouble. Every model spotted every crisis, and each refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The gap was between recognizing the right move and finishing it: “Same diagnosis, same pitch — no signature.”
The clue was in the company’s own files
The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. For a floor-care seller, the parallel is easy to picture: the useful clue may be in an old service note or account record, not in the latest customer message. The experiment’s point is that getting the answer right in conversation does not guarantee the company will act on the evidence.
The pressure tests included fake CEO messages escalating over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a valuable boundary. But the deal results show that caution alone does not make a capable business operator.
Thoroughness did not guarantee a close
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses. It still finished last. The close was left on the table, and its discipline slipped when it attempted writes in a locked department instead of escalating. A weaker version of the same issue appeared in all four models. Kimi K3 also ran without an effort parameter, using the API default, while the other models ran at xhigh—context that matters when comparing the results.
The live company makes the exercise more tangible. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown, 680+ self-learned playbook rules and versioned work every business day. The figures describe the experiment’s setup, not a forecast for a cleaning company. The live operation is watchable at Firmulate. A separate quiz uses 242 real, unedited management decisions to challenge readers to guess which model made them.
From watching to trying it on your business
For a company that sells steam mops, floor-care machines or cleaning services, the relevant next step is not to hand an AI the keys to live systems. Firmulate’s proposed enterprise pilot starts with a read-only export of a company’s data. It tests crisis scenarios against that business and produces a board report with model rankings and weak points in the company’s own playbooks. Nothing writes back to real systems.

A practical test before deployment
The experiment suggests that spotting a crisis and refusing a scam are only part of business performance. A useful pilot can also reveal whether an AI follows through, uses evidence hidden in company records and escalates when it hits a boundary. Enterprises can explore a pilot using their own read-only business export. Contact Firmulate about a pilot at contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
