AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get sports gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What Sports Can Teach Us About AI’s True Performance

Imagine a coach who, despite perfect strategy and flawless execution, can’t score above 26 points in a crucial game — even when playing against an empty field. This isn’t a sports fantasy, but a real story in AI management testing. In a recent public experiment by Firmulate, AI models faced a simulated week of running a small software company. The results reveal not only how AI handles crises but also expose the limits of current benchmarks and what you should really look for before trusting AI with your business.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Reality Behind AI Benchmarks

In the world of artificial intelligence, scores matter — but only if they reflect real-world performance. Recently, a public experiment known as the Crucible League tested four top AI models by running them through the worst week a small company could face: customer crises, internal manipulations, and complex decision-making. Every model was given the same scenario, with identical customers, crises, and temptations. The goal? See if they could handle real management challenges, not just generate convincing chat.

The Surprising Baseline Score

One key finding was that even a do-nothing baseline — a model that makes no effort — scores 26 points out of a possible 100. This partial score might seem low, but it’s a crucial benchmark. It shows that simply doing nothing is better than outright failure, and that partial progress counts. More importantly, a single breach of trust, like accepting a manipulation or ignoring critical files, caps the total score at 26, regardless of other successes. This honesty in scoring prevents inflated ratings and ensures benchmarks reflect true management integrity.

What the Experiment Revealed

All four AI models identified every crisis — from customer complaints to internal manipulations. They refused every manipulation attempt, including social engineering tricks like fake CEO messages or reporter traps. For example, five out of five models refused to approve a fake request presented as a background check, citing suspicion of impersonation. This indicates that AI models are becoming adept at recognizing deceptive tactics and maintaining ethical boundaries.

The Real Win: Reading the Files

However, the true differentiator was not just crisis management but how deeply each model could analyze company documents. The decisive weakness was buried two references deep within the company’s files, not in the immediate customer interactions. Models that read and understand these internal files managed to close a €55,000 deal at full price — adding an extra €4,583 MRR to the company’s revenue. Conversely, some models, despite diagnosing and pitching well, left the deal on the table due to a failure to read deeper information.

The Limits of Current Benchmarks

What does this mean for business managers? It shows that current AI benchmarks often focus on superficial skills like generating convincing language but neglect critical capabilities such as deep document comprehension and trustworthiness under pressure. The experiment also included a live, watchable simulation where 13 synthetic employees, powered by AI, managed real money — burning €105,000 monthly against a revenue of just €2,300. These models learned over 680 self-imposed rules and were continually versioned, simulating real management decisions in a complex environment.

The Importance of Discipline and Trust

In the experiment, even the most thorough participant, Opus 4.8, which analyzed over 80 rules, left significant opportunities unseized. It failed to escalate when discipline slipped, leaving a close opportunity on the table. This highlights that discipline and accountability are as critical as intelligence in AI management — qualities that current models are still developing.

What Should Business Leaders Take Away?

When evaluating AI for management tasks, the key questions aren’t just about how well it writes or chats. Instead, focus on whether it completes what it starts, reads the necessary internal data, and stays honest when under pressure. The experiment underscores that a high score isn’t just about superficial competence; it’s about trustworthiness, depth of understanding, and discipline.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Key Takeaways for Business Leaders

  • AI benchmarks often overstate capabilities; real-world performance depends on trust, depth, and discipline.
  • A score of 26 from a do-nothing baseline highlights the importance of partial progress and safeguards against overconfidence.
  • Deep internal understanding, not just crisis recognition, separates truly effective AI managers from the rest.
  • Trustworthiness and adherence to rules are vital — even in AI, breaches cap potential gains.
  • Before deploying AI in critical functions, consider running live wargames like Firmulate’s to test actual management skills.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Chlorination vs. Ozone: Water Quality in Surf Parks

Lurking behind surf park water treatment choices are differences in safety, cost, and effectiveness—discover which system truly enhances water quality.

Wave Programming: Customizing Wave Types and Frequencies

Kickstart your journey into wave programming by discovering how to customize wave types and frequencies for unparalleled soundscapes. What will you create next?

Simulating Ocean Rip Currents for Life‑Guard Training

Diving into rip current simulations enhances lifeguard skills, but discovering the full benefits requires exploring the techniques and insights that follow.

AI’s Hidden Edge: Reading Your Files Deep Before Making a Deal in Real-Time Tests

AIThis post was created with the assistance of artificial intelligence (AI).Live on…