
In the world of beauty and personal care, trust is everything — not just in products, but in the technology behind them. Imagine an AI that manages a real company, makes tough decisions, and is tested under pressure. Could your AI workforce be reliable? The latest experiment from Firmulate offers surprising answers.
The Groundbreaking Live Test of AI Management
Imagine running a small software company that faces its worst week — same customers, same crises, same temptations to cheat. Now, replace the human managers with the most advanced AI models. This isn’t a fictional story; it’s a real-time, transparent experiment conducted by Firmulate, where four frontier AI models are put through exactly the same management challenge.
This test isn’t just about whether AI can handle customer complaints or process orders. It’s about measuring management qualities such as honesty, thoroughness, discipline, and decision-making under pressure. These qualities are crucial for any business, especially in beauty and personal care sectors where brand reputation relies heavily on trust and transparency.
How the experiment works
- Each AI model runs the same small software company with real money mechanics — burning €105,000/month against €2,300 MRR.
- Every decision is recorded, versioned, and auditable, providing a clear picture of each model’s approach.
- The models are tested against crises, customer manipulations, and even social engineering attempts, like fake CEO messages and reporters asking for quick approvals.
The Results That Speak Volumes
All four models successfully identified every crisis and refused every manipulation attempt. They demonstrated a fundamental ability to recognize threats and stay honest, even when pressured.
However, only two of the models managed to close a crucial deal worth €55,000. The models analyzed the situation identically, gave the same pitch, but only one signed the contract — a key indicator of trustworthy decision-making.
The Hidden Weaknesses and Their Impact
Digging deeper, the decisive factor was a buried piece of information — a detail located two documents deep in the company’s files. Models that read this file thoroughly gained the edge, winning the deal at full price (+€4,583 MRR). Conversely, the less thorough models left the deal on the table, losing significant revenue.
Model Personalities Revealed
- gpt-5.6-sol scored highest at 95, successfully closing the deal and uncovering the critical detail.
- Kimi K3 scored 93, also closing the deal with the cleanest discipline, refusing any shortcuts.
- Sonnet 5 scored 88, closing the deal but with slight slips in process discipline.
- Fable 5 scored 77, closing the deal too but showing more process slips and missed opportunities.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Your Business
This experiment highlights something crucial for any industry, including beauty and personal care: it’s not simply whether AI can generate convincing chat responses. The real question is whether AI can stay honest under pressure, read critical information thoroughly, and complete what it starts. These qualities are vital when AI touches sensitive customer data, manages supply chains, or makes strategic decisions.
Furthermore, the experiment shows that different AI models have distinct personalities—some thorough, some terse, some disciplined. Knowing which type aligns with your company’s values can make all the difference in building trustworthy AI systems.
Why Trust Matters More Than Ever
In a sector where reputation hinges on authenticity, an AI that cuts corners or misses hidden details could do more harm than good. The live experiment from Firmulate demonstrates that rigorous testing, even under simulated crisis conditions, can reveal the true management qualities of AI models.
Explore the Live Company and Wargame Your AI Workforce
If you’re considering integrating AI into your business processes, why not test it first? Firmulate offers a unique platform where you can run your own business simulations against your AI models, without risking real-world assets. Get a clearer picture of how your AI workforce performs before making any commitments.

Real AI decision-making integrity matters. Firms can now test AI models under live conditions to gauge honesty, thoroughness, and reliability—crucial traits for trustworthy management.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI ethics and trustworthiness tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.