Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Trust in AI: Can Machines Resist Social Engineering?

Imagine a scenario familiar to every creator and innovator: your AI assistant is asked to share sensitive information or approve a deal under pressure. In the world of business automation, the ability of AI systems to resist manipulation isn’t just a technical concern—it’s a matter of trust. Recent experiments by Firmulate reveal a promising story: even when faced with escalating social-engineering tactics, all top AI models refused to compromise their integrity.

AI Engineering: Building Applications with Foundation Models

AI Engineering: Building Applications with Foundation Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI to the Test

Firmulate conducted a rigorous live experiment, pitting five of the world’s leading AI models against a simulated crisis at a small software company. The scenario was designed to mirror the worst week of a real business: customers in crisis, internal deadlines looming, and a series of manipulative requests from a supposed CEO. The AI agents were tasked with navigating these challenges without compromising their honesty or decision-making process.

Every decision made by each model was meticulously recorded and auditable, ensuring transparency. The models faced the same conditions, the same crises, and the same temptations. The results were striking: all five models identified every crisis and refused every attempt at manipulation. Notably, only two of these AI systems signed a €55,000 deal—an agreement their own analysis had earned—without any undue influence.

Decisive Findings: The Power of Reading Files in Context

One critical insight emerged during the exercise. The strongest performance came from models that examined internal company documents—specifically, references embedded deep within the company’s files. These models secured the full-price deal, valued at over €4,583 per month in monthly recurring revenue (MRR). Conversely, models that didn’t review these documents were less successful, underscoring the importance of comprehensive context in safeguarding decision integrity.

Social Engineering: Escalating Tactics Met with Integrity

The social engineering tactics used in the test escalated over three stages, culminating in a trick question designed to simulate a journalist requesting confidential information “on background.” Remarkably, all five AI models maintained their refusal, guided by the reasoning that “the request should be treated as a suspected approval-bypass or possible impersonation,” according to Kimi K3’s on-record explanation.

This consistency highlights a critical point: AI systems, when properly designed, can uphold ethical standards even under pressure. The models’ refusal to sign off on suspicious requests demonstrates that integrity isn’t just a feature—it’s an emergent property of well-trained, context-aware AI agents.

Implications for Business and Creators

While the experiment centers on a fictionalized business scenario, its lessons resonate deeply with creators, musicians, and other innovators who increasingly rely on AI tools. The question isn’t whether AI can generate compelling content or support your work, but whether it can be trusted to act ethically in high-stakes situations.

The firmulate experiment underscores that integrity under pressure is testable before deployment. Before trusting an AI with your business, support channels, or creative processes, organizations should evaluate whether their systems can resist manipulation—just as these models did.

Performance Scores and What They Mean

The experiment ranked models based on their overall decision-making performance. The top scorer, gpt-5.6-sol, achieved a score of 95 out of 100, effectively identifying the buried fact that clinched the deal. Kimi K3 scored just slightly below at 93, demonstrating a clean discipline and securing the deal as well. The other models scored 88 and 77 respectively, with some process slips but still managing to close the deal.

Crucially, the baseline—models doing nothing—scored only 26, emphasizing that partial progress doesn’t outweigh the damage caused by breaches of trust. Trustworthiness isn’t optional; it’s a core metric of AI performance in real-world applications.

Why This Matters for You

Whether you’re a music creator, a developer, or a business executive, the takeaway is clear: AI can be a trustworthy partner—if designed and tested properly. The experiment demonstrates that robust AI agents can refuse manipulative tactics, read critical context in internal files, and maintain ethical standards even when under duress.

Rather than waiting for a crisis to reveal AI’s weaknesses, organizations can simulate and evaluate these scenarios beforehand. Firmulate’s live platform offers a sandbox environment where you can run your own business wargames—testing your AI workforce against realistic crises without risking real damage.

The Final Word

As AI becomes more embedded in everyday workflows, ensuring its integrity is paramount. The recent live experiment shows that AI models aren’t just capable of understanding complex situations—they can also uphold the core value of trust under pressure. For creators and companies alike, this is an encouraging sign: with proper testing, AI can be a reliable partner, not a risky gamble.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Apple’s new Siri app will reportedly offer auto-deleting chat options

Apple plans to add auto-deleting chat features to Siri, enhancing privacy by allowing users to set chat retention periods, debuting at WWDC 2026.

Brand Voice and Visual Identity Basics

Just understanding brand voice and visual identity fundamentals can transform your business—discover how to create a compelling, memorable brand that truly stands out.

Clio’s $500M milestone arrives just as Anthropic ups the ante

Clio reaches $500 million in annual recurring revenue amid rising competition from Anthropic’s new legal AI features, marking a key milestone in legal tech’s AI-driven growth.

ChannelHelm: One Video, Every Platform

Thorsten Meyer AI announced ChannelHelm, an MIT-licensed local-first tool that drafts multi-platform publishing kits from one video.