
Imagine having a sous-chef who, instead of just tasting your dish, reads every ingredient list, every prep note, and every secret recipe before helping you cook. Would they make better decisions? Or could missing a hidden ingredient cost you the deal? In the world of AI, the ability to ‘read your files before answering’ is proving to be a game-changer—distinguishing those that succeed from those who fall short, especially in high-stakes scenarios.
The Experiment That Tested AI on a Small Company’s Worst Week
Recently, a groundbreaking live experiment put four leading AI models through the ultimate management test. Each AI was tasked with running a simulated small software company facing its worst week—same customers, same crises, and the same temptations to cheat. This wasn’t a simple chat test; every decision was carefully versioned, auditable, and rooted in realistic mechanics. The goal? See which AI could truly read between the lines and act honestly.
As an affiliate, we earn on qualifying purchases.
The Surprising Results: Reading Deep Wins
The outcome? All four models identified every crisis and refused manipulation attempts. Yet, only two managed to close a critical €55,000 deal, earning their own analysis. These models didn’t just diagnose problems—they understood the full context, including hidden details buried two document references deep within the company’s own files. Essentially, the winners read the memo that contained the key fact, giving them an edge that wasn’t visible in typical chat demos.
enterprise AI data comprehension tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Deep Reading Matters in AI Decision-Making
In real business scenarios, crucial information can be tucked away in internal documents, beyond the immediate customer interactions. If your AI only responds based on surface-level data, it might miss the finer details that could clinch a deal or prevent a disaster. The experiment underscores a vital truth: successful AI agents need to read and understand your files thoroughly before acting. Companies that invest in models capable of deep comprehension could see their AI become not just a chatbot but a strategic partner that reads the fine print.
As an affiliate, we earn on qualifying purchases.
Deception and Integrity Under Pressure
Another key aspect of the experiment tested social engineering—the classic fake CEO messages, escalating threats, and reporter tricks. All five models refused to be manipulated, demonstrating integrity. For example, Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline is crucial when AI is involved in sensitive decisions, from customer onboarding to financial approvals.
AI decision-making tools for businesses
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Company: Real Money, Real Risks
The experiment isn’t just theoretical; it’s set in a live simulation with a real-money company. Thirteen synthetic employees manage operations, burning €105,000 monthly against a modest €2,300 monthly recurring revenue (MRR). Every workday, the setup is versioned, and the live process is publicly watchable on firmulate.com/live. This ongoing test demonstrates how AI decision-making impacts real-world businesses, emphasizing the importance of reading and understanding internal data before acting.
What the Benchmarks Reveal
- GPT-5.6-sol scored 95, the highest, successfully finding the hidden fact and closing the deal.
- Kimi K3 scored 93, also closing the deal with excellent discipline.
- Sonnet 5 scored 88, and another Sonnet 5 scored 77, both closing the deal but with more slips in process discipline.
- The baseline score was just 26, highlighting how partial progress and breaches of trust are problematic.
Implications for Business and AI Adoption
The takeaway is clear: if AI is to be trusted with your CRM, support, or forecasting, it must do more than produce coherent chatter. It must read your internal files thoroughly, stay honest under pressure, and finish what it starts. Companies that incorporate models with these capabilities can potentially save or earn millions, as demonstrated by the €4,583 monthly recurring revenue tied to the successful deal.
Try It Yourself
Curious how your own AI workforce might perform? Firms can run the same management wargame against a read-only export of their business. It’s a safe, no-risk way to test whether your AI can identify hidden details and uphold integrity—before deploying it in the real world.
More details and live results are available at firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html