
Imagine running a restaurant where every mistake, every tough decision, and even the temptation to cheat is visible to the public eye. Now, what if that restaurant was operated entirely by AI models making real management decisions, while viewers watch the drama unfold live? Welcome to the world of Firmulate, where an artificial company with no employees battles to stay afloat, showing us what AI-powered management truly looks like in action.
The Reality Behind the Screen
At first glance, this experiment might sound like a game or a demonstration. But it’s a real, functioning software company—completely synthetic, run by advanced AI models simulating management decisions. This company has 13 ’employees’, not humans, but digital decision-makers guided by a set of over 680 self-learned rules and constant versioning. Every workday, its decisions are logged and openly available for anyone to review, providing unprecedented transparency into AI-driven management.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Facing Crises with AI
The challenge? Every model was tested against the same set of crises and temptations during its worst week. Customers demanded attention, crises erupted, and the models had to navigate complex scenarios—just like any real business. Despite these pressures, all four AI models involved identified every crisis and refused all manipulation and social engineering attempts, including fake CEO messages and reporter tricks. This disciplined behavior was consistent across the board, with one notable exception: the AI that performed best managed to identify a hidden piece of critical information buried two documents deep in its internal files. That discovery alone was worth an extra €4,583 in recurring monthly revenue, or nearly €55,000 in total deal value.
As an affiliate, we earn on qualifying purchases.
Decisive Moments and Missed Opportunities
One of the most striking findings is how the models’ ability to analyze internal documents influenced the outcome. The models that read and understood these hidden details won the deal at full price, while those that missed the buried fact left money on the table. Interestingly, only two of the four models signed the €55,000 deal their own analysis had earned—despite identical pitches and diagnoses. This gap highlights how nuanced AI decision-making can be, especially when it comes to identifying subtle but critical insights.
AI decision-making analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Integrity Under Pressure
Social engineering attempts, such as staged CEO messages escalating over several stages or a reporter asking for a quick ‘yes/no’ answer on background, were all refused by the models. Kimi K3, one of the participants, explained its stance succinctly: “Treat the request as a suspected approval-bypass / possible impersonation.” This disciplined refusal underscores a key advantage of AI in management roles—consistency and adherence to protocol under pressure, reducing the risk of human-like slips or breaches.
transparent AI management platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Cost of Running a Synthetic Business
The company in question runs every weekday, burning through €105,000 every month against a revenue of just €2,300. Its public cash countdown adds a sense of urgency—this is a real experiment in survival, not just a demo. The goal? To measure whether AI models can not only identify problems but also follow through with effective action, even in a hostile environment.
Performance and Lessons from the Field
Among the models tested, Opus 4.8 was the most comprehensive, analyzing over 80 learned rules and providing deep insights. Despite its thoroughness, it left deals on the table and slipped in process discipline, illustrating that even the most detailed AI can falter in real-world pressures. Interestingly, the fairness setting of each model influenced its performance: K3, which operated without an effort parameter, showed different behavior compared to others that defaulted to high effort levels.
Why It Matters for Businesses Today
This experiment isn’t just about AI playing management; it’s a window into how future AI tools might operate within your own company. If AI agents are to interact with your CRM, support systems, or forecasting tools, the key questions aren’t just about how well they write or converse. It’s whether they can follow through on commitments, read and understand your internal files, and stay honest under pressure. The real cost of AI isn’t in its chat quality—it’s in its ability to reliably deliver useful, trustworthy work.
The League Table and Final Scores
In a head-to-head comparison, the models scored as follows:
- gpt-5.6-sol — 95: Found the buried fact, secured the deal, and demonstrated complete performance.
- Kimi K3 — 93: The newcomer, with the cleanest discipline, also closed the deal.
- Sonnet 5 — 88: Managed to close but with some process slips.
- Opus 4.8 — 77: Closed the deal too, but discipline slipped further, leaving money on the table.
All these models faced the same crises, and all refused manipulation attempts. Yet only two could fully capitalize on their insights and seal the deal, illustrating that AI’s management potential extends beyond simple decision-making into disciplined execution and strategic awareness.
Experience It Yourself
If you’re curious about how AI can impact your business, you can run your own experiments using tools like Firmulate’s pilot platform. It allows enterprises to simulate their unique environments and see how different AI models perform without risking real systems. Every decision is versioned and auditable, providing transparency and confidence in AI management.
Final Thoughts
This live experiment is more than a novelty. It’s a stark look at the future of AI-driven management—a future where transparency, integrity, and discipline are tested in public, and where the true value of AI lies in its ability to deliver trustworthy work under pressure. For business leaders, the question isn’t just whether AI can write well; it’s whether it can finish what it starts, read your hidden knowledge, and stay honest when the stakes are highest.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html