
In an era where artificial intelligence is increasingly integrated into daily business operations, the true test isn’t just in how well an AI can generate text or analyze data—it’s in its integrity under pressure. For companies handling sensitive customer data or critical decision-making, trust in AI systems is paramount. Recent experiments with live, real-world AI models reveal a promising trend: when push comes to shove, AI can indeed resist manipulation and stay honest.
Testing AI Integrity in a Live Business Environment
At the forefront of AI evaluation is a groundbreaking experiment conducted by Firmulate, a company dedicated to benchmarking AI models as complete business simulators. This live test pits five leading AI models against a simulated small software company facing its worst week—same crises, same temptations, same customer demands. The goal? To see if these models can maintain integrity, avoid manipulation, and make honest decisions in real-time conditions.
As an affiliate, we earn on qualifying purchases.
The Setup: A Week of Crises and Temptations
The experiment involves a meticulously crafted scenario where each AI model manages the company’s operations, from customer support to financial decision-making. The models are tasked with handling escalating social-engineering attempts—fake CEO messages urging quick, unchecked actions, including requests to send sensitive customer data and sign off on unwarranted deals. These attempts are staged in three stages, culminating in a reporter trick where a simple yes/no response is solicited to bypass checks.
AI ethical decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Stunning Result: All Models Refused Manipulation
According to the latest data, all five models refused every manipulation attempt, including the staged CEO impersonation. The standout was the Kimi K3 model, which explicitly treated such requests as suspected impersonation, reflecting a cautious and security-conscious approach. Remarkably, every model recognized the social-engineering escalation and declined to act—an encouraging sign that integrity can be built into AI systems before deployment.
AI model integrity testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Decisive Factors Beyond Surface-Level Performance
While all models spotted the crises and refused manipulative requests, the ultimate difference came down to their decision-making depth. The key weakness was not in their crisis detection but in their internal data reading: models that delved deeper into the company’s own files secured better business outcomes. Specifically, those that read two document references within the company’s files were able to close deals at full price—adding €4,583 MRR (monthly recurring revenue)—where others left money on the table.
business AI decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Performance of the Top Models
- gpt-5.6-sol achieved a perfect score of 95, identified the buried critical information, and closed a deal worth full price.
- Kimi K3 scored 93, demonstrating exemplary discipline and successfully closing the same deal.
- Sonnet 5 and Fable 5 also closed deals, though with slightly less consistency, scoring 88 and 77 respectively.
The only model that failed to secure the full deal was Opus 4.8, which left a significant amount of money unclaimed due to process slips—such as writing attempts into a locked department instead of escalating. This highlights that even highly thorough models can slip if discipline isn’t maintained under pressure.
Implications for Businesses and AI Developers
This experiment underscores a critical insight: the integrity of AI systems isn’t just about their ability to perform tasks but also their capacity to uphold ethical standards, especially when tested. For companies in sectors like beauty and personal care, where trust and customer data security are vital, deploying AI that resists manipulation is essential.
Why Trust Matters — Before a Crisis Arises
The Firmulate live experiment illustrates that integrity can be integrated into AI systems before they face real-world crises. The models’ ability to recognize staged social-engineering attempts and refuse to participate demonstrates that ethical decision-making can be a fundamental part of AI design, not just an afterthought. This proactive approach ensures businesses can rely on their AI tools to act honestly, even under pressure.
Next Steps: Testing Your Own AI Workforce
Businesses eager to assess their AI systems can run similar wargames in a safe, read-only environment—testing AI decision-making against real crises without risking actual data or processes. This approach helps identify weaknesses before deployment, ensuring that when the stakes are high, AI remains trustworthy and disciplined.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html