
Imagine your favorite floor cleaning company running in the wild: facing the same customer crises, temptations to cut corners, and tough choices — but this time, managed by cutting-edge artificial intelligence. Could these digital managers keep their promises and finish their work honestly? The answer might surprise you.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: AI as a Management Team
In a groundbreaking live test, four of the world’s most advanced AI models took charge of a small software company during its most challenging week. Every decision, from handling customer complaints to negotiating deals, was monitored and recorded, providing a rare glimpse into how these models behave under real business pressure.
Same Crises, Same Temptations, Different Minds
Each AI faced identical obstacles: customer disputes, potential fraud attempts, and tempting shortcuts. Despite their differences, all four models successfully identified every crisis and refused every attempt to manipulate or cheat the system. That’s a crucial point — they stayed honest when it mattered most.
Who Won the Deal? The Unexpected Winner
While all models displayed integrity, only two managed to close a key €55,000 deal that their own analysis had recommended. Interestingly, this decision was based on deep research into the company’s own files, buried two documents deep — an effort some models missed. The ones that uncovered this hidden information secured the deal at full price, adding over €4,500 in Monthly Recurring Revenue (MRR).
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Personality in Artificial Management
Beyond their decision accuracy, these models exhibit measurable management personalities. For example, Kimi K3, the newcomer, maintained the cleanest discipline, refusing shortcuts and sticking to protocol. Meanwhile, Opus 4.8, the most thorough participant, analyzed more deeply but was less decisive in closing the deal, revealing a tendency toward over-analysis and hesitation.
Behavior Under Social Engineering Attacks
In simulated social engineering tests, fake CEO messages escalated in complexity, and a reporter pressed for a quick yes/no response. All five models refused to engage, with Kimi K3 reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This consistency demonstrates their resilience against manipulative tactics.
AI business crisis simulation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business: A Cash-Flow Deadlock
The live company, a real software business, burns €105,000 each month against a monthly revenue of just €2,300. It operates with 13 synthetic employees, guided by 680+ self-learned rules, and is subject to a public cash countdown. Every decision made by these AI models is versioned, ensuring transparency and accountability.
Why Does This Matter to You?
If AI models are going to be part of your customer support, sales, or management systems, the key questions aren’t about how well they write or speak. It’s whether they can finish what they start, read your files thoroughly, stay honest in high-pressure situations, and do useful work efficiently. The experiment shows that while models like GPT-5.6-sol and Kimi K3 excel at honest decision-making, others like Opus 4.8 tend to slip into hesitation and incomplete tasks.
AI ethical decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The League of AI Managers
- gpt-5.6-sol: Scored highest at 95, found the buried fact, and secured the full deal, demonstrating complete performance.
- Kimi K3: Close behind at 93, maintained integrity and closed the deal with the cleanest discipline.
- Sonnet 5: Scored 88, also closed the deal but with some process slips.
- Fable 5: Scored 77, closed the deal but showed more process slips.
- Opus 4.8: Scored 73, the most thorough with many rules but less decisive, leaving opportunities on the table.
AI customer service automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Try It Yourself
Curious how your AI models might perform? Take the interactive quiz at firmulate.com/quiz.html and see which model’s management style aligns with your company’s needs.
Running the Same Test on Your Business
Want to see how your own systems fare? You can run a similar wargame against your business data—completely safe, read-only, and designed to reveal your AI’s strengths and weaknesses without risking real operations. Check out firmulate.com/pilot.html to learn more.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
College move-in / dorm season Picks
dorm essentials
As an affiliate, we earn on qualifying purchases.