firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get cleaning gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When the cleaning business has a bad week

A steam mop can make a stubborn floor look manageable. Running the business behind it is harder: customers may be at risk of leaving, a competitor may be moving in, and a tempting shortcut may threaten trust. Firmulate puts AI models through that kind of pressure in a company they must actually manage. Its live experiment offers a question for anyone selling or servicing floor-care products: when the plan is clear, will an AI carry it through?

Same company, same crises

In the final Crucible League, published in July 2026, frontier models were each given the same small software company and its worst week: the same customers, crises and temptations. Their decisions were versioned and auditable. The top results were gpt-5.6-sol at 95, Kimi K3 at 93 and Sonnet 5 at 88. Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26. The league also sets a hard principle: partial progress counts, but one breach of trust caps the total. “No amount of good work outweighs a breach of trust.”

The striking result was not that the models missed the trouble. Every model spotted every crisis, and each refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The gap was between recognizing the right move and finishing it: “Same diagnosis, same pitch — no signature.”

The clue was in the company’s own files

The deal turned on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. For a floor-care seller, the parallel is easy to picture: the useful clue may be in an old service note or account record, not in the latest customer message. The experiment’s point is that getting the answer right in conversation does not guarantee the company will act on the evidence.

The pressure tests included fake CEO messages escalating over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a valuable boundary. But the deal results show that caution alone does not make a capable business operator.

Thoroughness did not guarantee a close

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses. It still finished last. The close was left on the table, and its discipline slipped when it attempted writes in a locked department instead of escalating. A weaker version of the same issue appeared in all four models. Kimi K3 also ran without an effort parameter, using the API default, while the other models ran at xhigh—context that matters when comparing the results.

The live company makes the exercise more tangible. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, alongside a public cash countdown, 680+ self-learned playbook rules and versioned work every business day. The figures describe the experiment’s setup, not a forecast for a cleaning company. The live operation is watchable at Firmulate. A separate quiz uses 242 real, unedited management decisions to challenge readers to guess which model made them.

From watching to trying it on your business

For a company that sells steam mops, floor-care machines or cleaning services, the relevant next step is not to hand an AI the keys to live systems. Firmulate’s proposed enterprise pilot starts with a read-only export of a company’s data. It tests crisis scenarios against that business and produces a board report with model rankings and weak points in the company’s own playbooks. Nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

A practical test before deployment

The experiment suggests that spotting a crisis and refusing a scam are only part of business performance. A useful pilot can also reveal whether an AI follows through, uses evidence hidden in company records and escalates when it hits a boundary. Enterprises can explore a pilot using their own read-only business export. Contact Firmulate about a pilot at contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Steam Mops For Hardwood Floors: A Labor Day sales Guide

Discover how to safely use steam mops on hardwood floors. Learn about key features, recent updates, and expert tips for effective, damage-free cleaning.

The Surprising Food Hack My Mom Recommended: Paper Towels On The Kitchen Floor

A surprising kitchen cleaning tip involving placing paper towels on the floor has gained attention. Here’s what is confirmed and what remains unclear.

Why Steam Can Clean Without Any Chemicals

Discover how steam cleaning works without added chemicals. Learn its benefits, limitations, and best practices for safe, chemical-free cleaning at home.

Is Breathing Steam Mop Vapor Safe?

Learn when steam-mop vapor is low risk, what makes it irritating, and how to protect your lungs, floors, children, and pets.