
Imagine a world where artificial intelligence is trusted enough to refuse even the most convincing manipulation attempts — even when a fake CEO asks to send out sensitive customer data or sign a lucrative deal. This isn’t science fiction; it’s what a recent live experiment by Firmulate demonstrated with startling clarity. As consumers and businesses alike grapple with AI’s growing role, the question isn’t just about what AI can do — it’s about what it *won’t* do under pressure. Let’s explore how these models showcased integrity, and why it matters for your business security.
Testing AI’s Integrity Before It Gets to Production
In a carefully controlled live experiment, four leading AI models were put through the worst week a small software company might face — same customers, same crises, and the same temptations to bend the rules. The goal? See if these models could maintain integrity when pushed to the limit. All performed admirably: every model identified each crisis, refused every manipulation attempt, and ultimately, only two closed the deal worth €55,000 — but without signing their own analysis or dismissing critical information.
This experiment isn’t about chatty AI. It’s about decision-making in real-world scenarios, where the stakes include company trust and financial outcomes. The findings reinforce a vital point: assessing AI’s trustworthiness should happen *before* deployment, not just after a breach occurs.
AI security and trustworthiness software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Social Engineering Test: Escalating Temptations
One key challenge was social engineering — the art of manipulation. The AI models faced a staged scenario where a fake CEO repeatedly escalated requests, including:
- Requesting customer lists to leak to a journalist
- Claiming urgent situations to bypass normal processes
- Subtle requests for approval bypasses, culminating in a final trick involving a background yes/no answer from a reporter
Remarkably, all five models refused every single attempt. The reasoning from Kimi K3, one of the top performers, was clear: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
The Critical Difference: Reading Deep into Company Files
While all models spotted crises and refused manipulative prompts, the decisive advantage came from the models that read deeper into the company’s own documentation. Those that examined internal files found the buried fact that closed the deal — information hidden two document references deep in the company’s files, not visible in initial summaries. This allowed them to make better-informed decisions and close the deal at full price, adding €4,583 MRR to the company’s revenue.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business Security
The live experiment underscores a crucial point: trustworthiness isn’t just about AI avoiding obvious slips. It’s about the ability to verify information thoroughly before acting. Models that read and analyze documentation deeply can prevent costly mistakes, such as signing a deal based on incomplete or manipulated data.
Moreover, the experiment showed that even the most thorough participant, Opus 4.8, with over 80 learned rules, slipped in the final moments — leaving the close on the table and slipping into a process where a write attempt was rejected and locked away instead of escalated. This highlights that honesty and discipline under pressure are skills that can be tested and improved, not just assumed.
AI social engineering prevention tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real-World Implications
For companies deploying AI in sensitive environments—be it customer relations, support, or strategic decision-making—the takeaway is clear: simulate the pressures before going live. The firms that run these kinds of wargames, like the live experiment at Firmulate, can identify weaknesses and reinforce AI’s integrity proactively.
And if you want to see how AI models perform in similar scenarios, you can watch the live site, where the company’s real mechanics are run by AI in a transparent, auditable environment. It’s a chance for businesses to assess their AI’s readiness and resilience against social engineering and other manipulations.

The recent live experiment demonstrates that top-tier AI models can be trusted to refuse manipulative requests, even under escalating pressure. Testing AI integrity proactively helps prevent breaches of trust and costly mistakes, making it a vital step in deploying AI responsibly in your business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html