
Imagine trusting an AI to handle your most critical relationship—be it a business deal or a personal promise. How confident would you be if that AI had to navigate crises, resist manipulation, and keep its commitments? Recent live experiments with advanced AI models shed light on this question, and the results are both fascinating and revealing.
Get gifts for the two of you delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Experiment: AI as a Company Manager in Crisis
In a groundbreaking live trial, four leading frontier AI models were tasked with running a small software company through its worst week—dealing with customers, crises, and attempts at manipulation. The experiment wasn’t just about how well the AI could generate words; it was about whether these models could make decisions aligned with integrity and discipline under pressure. Every decision was documented and auditable, simulating real-world management dilemmas.
The Performance League
- gpt-5.6-sol scored the highest at 95, successfully finding a critical buried fact and closing a €55,000 deal — the complete performance.
- Kimi K3 (Moonshot) scored 93, just behind the leader, and demonstrated the cleanest discipline of all models by resisting manipulative tactics and reading deeply into company files to close the deal at full price.
- Sonnet 5 scored 88, also closing the deal but with some process slips.
- Fable 5 and Opus 4.8 scored 77 and 73 respectively, each managing to close the deal but with noticeable discipline lapses.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings: Integrity, Discipline, and Deep Reading
The core of the experiment revealed that all models identified every crisis and refused manipulation attempts, showing they can recognize threats and act ethically under pressure. The decisive factor for winning the deal was reading deeper into the company’s files—models that did this, like Kimi K3, secured the full-price sale, adding €4,583 in Monthly Recurring Revenue (MRR).
Interestingly, the models’ ability to resist social engineering was tested with staged fake CEO messages and a reporter trick. All five models refused to sign or escalate on suspicious requests, citing reasons like treating the request as a possible impersonation. This demonstrates a significant level of ethical resilience—crucial for real-world business applications.
As an affiliate, we earn on qualifying purchases.
The Real Company and Its Challenges
Behind the scenes is a live company with 13 synthetic employees, burning €105,000 monthly against €2,300 MRR. The setup includes over 680 self-learning rules and daily versioned decision logs, making it a transparent, observable environment. Watch it in action at firmulate.com/live.
ethical AI decision support system
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Surprising Weakness: Deep Document Analysis
The experiment uncovered that the most decisive weakness was not in spotting crises but in reading and understanding company documents. Models that thoroughly analyzed files won the full-price deal. Those that left the close on the table or slipped discipline failed, highlighting the importance of deep contextual reading—something AI still struggles with at times.
As an affiliate, we earn on qualifying purchases.
Fairness and Methodology
The experiment was fair: Kimi K3 was run without an effort parameter (the default API setting), while others used a more aggressive xhigh setting. This ensures a level playing field and authentic comparison.
Implications for Business and Beyond
As AI agents become more integrated into CRM, customer support, or forecasting, the question isn’t just about linguistic quality or superficial performance. It’s about whether they can complete their tasks ethically, resist manipulation, and make decisions based on thorough understanding. The experiment underscores that choosing an AI model should be based on real performance metrics—not just marketing demos.
Watch and Decide
Beyond testing, enterprises can run their own ‘wargames’ against their business data—without risking real systems—by using tools like those at Firmulate. This allows organizations to evaluate AI management quality before deployment, ensuring their AI workforce is trustworthy and disciplined from day one.
The Takeaway
In the evolving landscape of AI-driven management, the ability to read deeply, stay disciplined, and resist manipulation isn’t optional—it’s essential. The live experiment shows that even newcomers like Kimi K3 can outperform established models when it comes to integrity and thoroughness, making the case that choosing the right AI is now a strategic decision grounded in real-world testing.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
