
Imagine choosing a chef’s knife not just by its sharpness but by whether it can handle a busy night’s dinner rush without missing a beat. In the world of AI-driven management, the same principle holds. Newcomers can outperform established giants—not in talk, but in action. That’s the story emerging from a recent live experiment where AI models ran a real software company through its toughest week, and the results are revealing.
Get kitchen gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Live Wargame: Putting AI to the Business Test
In an unprecedented live experiment, four frontier AI models were tasked with managing a small, real-world software company during its most challenging week. This isn’t just a simulation; it’s a fully operational company with synthetic employees, real money mechanics, and daily decision-making. The goal was straightforward but demanding: see if these AI models can navigate crises, resist manipulative tactics, and ultimately close a crucial deal worth €55,000 in recurring revenue.
AI management software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Results: A Clear Leader Emerges
Scores from the experiment placed GPT-5.6-sol at the top with a score of 95, followed closely by the newcomer Kimi K3 at 93. The others—Sonnet 5, Fable 5, and Opus 4.8—scored 88, 77, and 73 respectively. Notably, all models identified every crisis and refused manipulation attempts, underscoring their integrity. Yet, only two models successfully closed the deal their own analysis had earned. Kimi K3, representing Moonshot, was among the victorious, demonstrating a clean discipline that outperformed older, more established models.
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness and the Winner’s Edge
Deep in the files—two document references into the company’s internal files—lie critical clues for closing the deal. Kimi K3 was the only model to uncover this buried data, giving it an advantage in decision-making. While all models recognized crises and resisted false requests, it was the ability to read and interpret hidden information that made the difference.
AI cybersecurity tools for social engineering
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Resisting Social Engineering and Manipulation
During the experiment, all AI models faced escalating fake CEO messages and a reporter trick—tests designed to manipulate decision-making. Remarkably, all five models refused these manipulations. Kimi K3’s reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation,” showing a cautious and disciplined approach that prevented breach of trust.
AI data analysis software for hidden insights
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Business: A Company Burning Cash
The company in the experiment operates with 13 synthetic employees and manages business mechanics that burn through €105,000 monthly against only €2,300 in monthly recurring revenue. This live setup is maintained daily, with over 680 self-learned rules guiding operations. Every decision is versioned and auditable, making it a transparent test bed for AI management capabilities. You can watch the company’s progress and decisions in real time at firmulate.com/live.
Insights and Implications for Business Leaders
The experiment reveals something critical: the true test of AI in management isn’t just whether it can generate convincing chat responses. It’s whether it can finish tasks, interpret hidden data, resist manipulative tactics, and stay disciplined under pressure. The fact that the newcomer Kimi K3 outperformed established models in these respects suggests that choosing an AI partner based solely on superficial benchmarks may be a mistake. The league table shows a clear open contest: the model that reads deeply, resists manipulation, and closes deals effectively is the one that matters.
Fairness and Methodology
It’s important to note that Kimi K3 ran without an effort parameter, using the API default setting, while all other models ran at a higher, xhigh effort level. This detail speaks to the efficiency of the K3 approach, making its performance even more impressive given its lower resource configuration.
What’s Next? Testing Your Business
For business leaders curious about AI’s potential, there’s a way to run similar tests on your own operations. Using Firmulate’s live platform, enterprises can simulate their own worst weeks—without risking real systems or data—by running a read-only export of their business. It’s a chance to see how your current management AI stacks up against the best and worst of the field, before making any commitments.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
