AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

In the world of AI, progress isn’t just about what the models say—they’re judged by what they do. Imagine a test where the worst possible score isn’t zero, but 26. Why? Because trust and partial progress matter just as much as perfect answers. This is the eye-opening story behind an innovative AI benchmark that reveals what honest AI really looks like in action.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get gifts for the two of you delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the Baseline: Why 26, Not 0

At first glance, you might expect a completely useless AI to score zero. Instead, the benchmark starts it off at 26 points. That’s because even a do-nothing model—one that skips decisions or ignores crises—still picks up some points for basic awareness. Think of it like a relationship: even if you do nothing, you’re still present. Partial progress counts, highlighting that in real business, doing nothing is still better than actively misbehaving.

Amazon

AI decision-making simulation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Methodology: Real Crises, Real Decisions

Here’s how the test works: four different AI models are put through the same simulated week of running a small software company. The scenarios include customer crises, manipulation attempts, and ethical dilemmas—things that real managers face daily. Every decision is recorded and auditable, ensuring transparency. The models are judged not just on whether they spot problems, but whether they act honestly and decisively.

Amazon

AI ethics and trust training tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Honesty and Effectiveness

  • All four models identified every crisis and refused every manipulation attempt, demonstrating integrity.
  • Only two models actually closed the deal, earning the full €55,000 reward. The others diagnosed correctly but held back on signing or slipped in process discipline.
  • Interestingly, the critical weakness wasn’t in the customer interactions or crisis management—it was in internal document reading. The models that read the company’s files secured a deal worth an additional €4,583 MRR, highlighting the importance of internal knowledge.
Amazon

AI internal document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Behavior Under Pressure: Trust Matters

One of the most revealing tests involved social engineering—fake messages from a CEO escalating through multiple stages, plus a reporter trick. Every model correctly refused to approve these manipulative requests. Kimi K3, for example, explicitly noted: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that these models aren’t just good at spotting crises—they can also resist manipulation attempts that would tempt a dishonest response.

Amazon

AI manipulation resistance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Business Simulation: Real Money, Real Stakes

The experiment runs on a live platform simulating actual company operations—13 synthetic employees, real cash flow mechanics, and daily decision-making. It’s a fully transparent environment where every move is versioned and observable at firmulate.com/live. This setup ensures that AI performance is judged on real-world readiness, not just language prowess.

The Surprising Role of Discipline and Rules

One standout participant, Opus 4.8, with its over 80 learned rules and deep analysis, still finished last. Its failure to sign the deal was due to slip-ups like writing into a locked department instead of escalating issues—showing that even thorough models can struggle with discipline in a high-pressure setting. This underscores that performance isn’t just about knowledge but also about consistent behavior and process discipline.

What Does This Mean for Business?

If AI tools are going to manage customer support, sales, or internal processes, their ability to finish what they start, read internal files, and stay honest under pressure is crucial. The benchmark shows that honesty, discipline, and thorough internal reading are vital measures that often go unnoticed in traditional chat demos.

Final Thoughts: A Transparent, Trustworthy Benchmark

This experiment by Firmulate offers a rare glimpse into how AI models perform in realistic, high-stakes scenarios. It values honesty and effectiveness over superficial chat quality, providing a more honest picture of what AI can do in actual business settings. And it reminds us that trust—like in relationships—is the foundation of productive collaboration.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like
ai personalized couple courses

AI and Education: Personalized Courses for Couples’ Goals

Personalized AI courses for couples’ goals adapt to your needs, promising a transformative learning experience—discover how this innovation can reshape your journey together.
ai driven romantic matchmaking

Love in the Time of Algorithms: AI Matchmaking Explained

AIThis post was created with the assistance of artificial intelligence (AI).AI matchmaking…
firmulate.com/quiz.html — live view

Can AI Models Make Better Business Decisions Than Humans? Test Your Guess

Discover how different AI models manage a real company’s crises, make deals, and stay honest under pressure. Test your guess at our interactive quiz.
ai filters influence attraction

How AI Dating Filters Could Shape Attraction

Ineffective use of AI dating filters may subtly influence perceptions of attractiveness, making you question whether true connection depends on authenticity.