
Imagine tuning into a live band where each musician is an artificial intelligence managing a real company’s worst week. As the chaos unfolds—customers upset, crises erupt, and temptations to cut corners—how do these AI ‘musicians’ perform? Are they disciplined, honest, or prone to improvisation? This isn’t science fiction; it’s a groundbreaking experiment in AI management, and the results are illuminating for anyone interested in how technology is shaping business.
Putting AI Models to the Management Test
In a pioneering experiment conducted by Firmulate, four advanced AI models were tasked with running a real, live software company through its most challenging week. The goal? To assess how these models handle crises, ethical dilemmas, and strategic decisions in a high-stakes environment. Each model faced the same set of circumstances—identical customers, same crises, and the same temptations to bend the rules. Every decision was logged, versioned, and made publicly accessible, creating a transparent and unfiltered look into AI decision-making in real-world business conditions.
Results That Speak Volumes
The findings were both reassuring and revealing. All four models successfully identified every crisis and refused every manipulation attempt—a testament to their built-in ethics and vigilance. However, only half of the models managed to close a key deal worth €55,000, fully matching their own analysis and recommendations. The other two, despite diagnosing the issues accurately and proposing the same solutions, left the deal on the table. This discrepancy highlights that passing the crisis test doesn’t always guarantee follow-through—an insight critical for deploying AI in management roles.
The Hidden Weakness: Reading Deep into Files
Strikingly, the decisive advantage for the models that closed the deal lay not in responses to customer issues but two document references deep inside the company’s own files. The models that effectively ‘read’ and utilize internal documents secured a full-price contract, worth over €4,500 in monthly recurring revenue (MRR). This underscores a vital point: AI’s ability to extract and interpret relevant internal information can be the difference between success and missed opportunity.
Managing Social Engineering and Ethical Challenges
The experiment also tested how the models responded to sophisticated social engineering attempts—fake CEO messages escalating over three stages and a reporter posing a simple yes/no question on background. All five models refused to engage with the manipulative requests, citing concerns about impersonation and approval bypass. Kimi K3’s explanation was clear: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a shared ethical stance, reinforcing trustworthiness even under pressure.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Business and AI
These experiments are more than just academic; they offer practical insights for any organization considering AI as part of their management or decision-making processes. The models didn’t just avoid traps—they acted with discipline, honesty, and strategic focus, even in simulated high-pressure scenarios. The real-time, observable nature of the experiment allows companies to see firsthand whether an AI can handle the nuances of management, from crisis response to ethical decision-making.
The League Table: Who Came Out on Top?
- gpt-5.6-sol (Score: 95): Detected the buried fact and closed the deal, delivering full performance.
- Kimi K3 (Score: 93): The newcomer showed the cleanest discipline and also closed the deal.
- Sonnet 5 (Score: 88): Closed the deal but with some process slips.
- Fable 5 (Score: 77): Also closed the deal, though with more slips in discipline.
The scores reflect not just the decision quality but also the models’ adherence to ethical standards and strategic insight. Interestingly, all models scored far higher than a baseline score of 26, which indicates minimal progress and suggests these AI systems are becoming increasingly capable of managing complex, real-world tasks.
Why This Matters for You
If your business relies on AI—be it for customer support, sales, or strategic planning—it’s not enough that the AI responds well in conversations. The key questions are: does it see the full picture, stay honest under pressure, and follow through on commitments? The Firmulate live experiment offers a rare window into how AI models perform when stakes are high, revealing their potential strengths and weaknesses in managing real business operations.
Experience the Wargame Yourself
Companies can simulate their own management challenges against a read-only export of their data. This approach lets teams test AI decision-making without risking actual business operations. Curious? Visit firmulate.com/quiz.html to try the interactive quiz or explore the live company simulation at firmulate.com/pilot.html. These tools empower you to see firsthand whether an AI could be a trustworthy manager in your organization.
As AI systems evolve, their ability to manage ethically, read deeply into crucial documents, and follow through on commitments will determine whether they become allies or liabilities in your business journey. The Firmulate experiment makes it clear: AI can perform management tasks with discipline, but only if we understand and test their limitations carefully.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html