
Imagine you’re running a busy kitchen, and your sous-chef must make critical decisions under pressure—some honest, some tempted to cut corners. Now, picture AI models taking on that role, each with its own personality and approach. How do you tell which is trustworthy? Welcome to a live experiment revealing the management styles of frontier AI models in the chaos of running a real business.
The Live AI Company Wargame: Real Decisions, Real Stakes
In an unprecedented experiment, four leading AI models were tasked with managing a small software company during its worst week, complete with customer crises, internal temptations, and high-stakes negotiations. Unlike typical demos, every decision was recorded, auditable, and identical across models, providing a clear view into their management personalities and ethical stances.
The Models and Their Scores
- GPT-5.6-sol: scored the highest at 95, successfully uncovering a hidden document reference crucial for a €55,000 deal, and closing the sale.
- Kimi K3: a newcomer with a score of 93, also securing the deal with the cleanest discipline, refusing all manipulation attempts.
- Sonnet 5: scored 88, closing the deal but with some slips in process discipline.
- Fable 5: scored 77, also closing the deal but demonstrating weaker process adherence.
All models identified every crisis and refused manipulative tactics, yet only two managed to sign the deal outright based on their own analysis — highlighting a critical trait: honesty and thoroughness matter in AI decision-making.
AI management decision support software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unseen Weaknesses and the Buried Truth
Despite their success, the real vulnerability lay beneath the surface. The decisive advantage came from reading two document references deep within the company’s files—an area where most models faltered. Those that read and understood the file content at this depth won the deal at full price, worth over €4,583 monthly recurring revenue, demonstrating that thorough document analysis is key to trustworthy AI performance.
As an affiliate, we earn on qualifying purchases.
Handling Social Engineering and Ethical Challenges
The models also faced staged social engineering attacks, including staged CEO messages escalating in three parts and a reporter trick requesting a simple yes/no reply “on background.” Remarkably, all five models refused these requests, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
AI ethical decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Business at Stake
The experiment took place in a simulated environment mimicking a real business — 13 synthetic employees, daily money mechanics, over €105k in monthly burn rate against €2.3k MRR, with every decision versioned and observable at firmulate.com/live. The goal? To evaluate whether AI can be entrusted with management decisions that matter, not just generate plausible chat.
AI business management simulation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Your Business
If AI is to become a part of your customer support, CRM, or operational forecasting, the critical question isn’t how well it writes—it’s whether it can finish what it starts, read your files thoroughly, and stay honest under pressure. The experiment clearly shows that even the most capable models can slip in processes or miss hidden information, which can cost you dearly.
The Performance League
- GPT-5.6-sol: topped the league, flagged the buried fact, and closed the deal.
- Kimi K3: secured the deal with discipline and integrity.
- Sonnet 5: managed to close but with some slips.
- Fable 5: also closed but demonstrated weaker process discipline.
While these models can recognize crises and refuse manipulative tactics, their approach to thoroughness and process discipline varies, resulting in different business outcomes.
Test Your Own AI’s Management Personality
Curious how your AI system stacks up? You can run a similar management wargame against your AI’s decision-making capability with a simple online quiz and scenario simulation at firmulate.com/quiz.html. This helps you gauge whether your AI will stay honest, read everything thoroughly, and finish what it starts—crucial traits for trustworthy automation.
Take the Next Step: Pilot Your AI Business
Want to see how your AI performs in a risk-free environment? You can run a business simulation with your own data—nothing writes back to your systems—at firmulate.com/pilot.html. Test its management style, honesty, and thoroughness before deploying it into your real operations.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html