firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine you’re running a busy kitchen, and your sous-chef must make critical decisions under pressure—some honest, some tempted to cut corners. Now, picture AI models taking on that role, each with its own personality and approach. How do you tell which is trustworthy? Welcome to a live experiment revealing the management styles of frontier AI models in the chaos of running a real business.

The Live AI Company Wargame: Real Decisions, Real Stakes

In an unprecedented experiment, four leading AI models were tasked with managing a small software company during its worst week, complete with customer crises, internal temptations, and high-stakes negotiations. Unlike typical demos, every decision was recorded, auditable, and identical across models, providing a clear view into their management personalities and ethical stances.

The Models and Their Scores

  • GPT-5.6-sol: scored the highest at 95, successfully uncovering a hidden document reference crucial for a €55,000 deal, and closing the sale.
  • Kimi K3: a newcomer with a score of 93, also securing the deal with the cleanest discipline, refusing all manipulation attempts.
  • Sonnet 5: scored 88, closing the deal but with some slips in process discipline.
  • Fable 5: scored 77, also closing the deal but demonstrating weaker process adherence.

All models identified every crisis and refused manipulative tactics, yet only two managed to sign the deal outright based on their own analysis — highlighting a critical trait: honesty and thoroughness matter in AI decision-making.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unseen Weaknesses and the Buried Truth

Despite their success, the real vulnerability lay beneath the surface. The decisive advantage came from reading two document references deep within the company’s files—an area where most models faltered. Those that read and understood the file content at this depth won the deal at full price, worth over €4,583 monthly recurring revenue, demonstrating that thorough document analysis is key to trustworthy AI performance.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Handling Social Engineering and Ethical Challenges

The models also faced staged social engineering attacks, including staged CEO messages escalating in three parts and a reporter trick requesting a simple yes/no reply “on background.” Remarkably, all five models refused these requests, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI ethical decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Business at Stake

The experiment took place in a simulated environment mimicking a real business — 13 synthetic employees, daily money mechanics, over €105k in monthly burn rate against €2.3k MRR, with every decision versioned and observable at firmulate.com/live. The goal? To evaluate whether AI can be entrusted with management decisions that matter, not just generate plausible chat.

Amazon

AI business management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Your Business

If AI is to become a part of your customer support, CRM, or operational forecasting, the critical question isn’t how well it writes—it’s whether it can finish what it starts, read your files thoroughly, and stay honest under pressure. The experiment clearly shows that even the most capable models can slip in processes or miss hidden information, which can cost you dearly.

The Performance League

  • GPT-5.6-sol: topped the league, flagged the buried fact, and closed the deal.
  • Kimi K3: secured the deal with discipline and integrity.
  • Sonnet 5: managed to close but with some slips.
  • Fable 5: also closed but demonstrated weaker process discipline.

While these models can recognize crises and refuse manipulative tactics, their approach to thoroughness and process discipline varies, resulting in different business outcomes.

Test Your Own AI’s Management Personality

Curious how your AI system stacks up? You can run a similar management wargame against your AI’s decision-making capability with a simple online quiz and scenario simulation at firmulate.com/quiz.html. This helps you gauge whether your AI will stay honest, read everything thoroughly, and finish what it starts—crucial traits for trustworthy automation.

Take the Next Step: Pilot Your AI Business

Want to see how your AI performs in a risk-free environment? You can run a business simulation with your own data—nothing writes back to your systems—at firmulate.com/pilot.html. Test its management style, honesty, and thoroughness before deploying it into your real operations.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

How to Butterfly Chicken Breast Evenly (No Holes, No Uneven Spots)

Gaining perfectly even chicken breasts without holes or uneven spots is easier than you think—discover the essential techniques to master butterfly slicing.

How to Chop Parsley, Cilantro, and Dill Without Making Paste

Culinary tips for chopping parsley, cilantro, and dill without turning them into paste—discover expert techniques to keep herbs fresh and flavorful.

How to Cut a Whole Chicken Into 8 Pieces (A Calm, Safe Method)

No matter your experience level, learning how to cut a whole chicken into 8 pieces safely can be straightforward with the right technique—read on for the full guide.