
When evaluating AI models for business, it’s tempting to assume that a tool doing nothing should score zero. But in a recent public benchmark experiment by Firmulate, even a passive AI baseline scored 26 points out of 100. That might seem odd — why doesn’t doing nothing get zero? And what does that say about how we measure AI readiness for real-world tasks?
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Surprising Benchmark Score of a Do-Nothing AI
In a live experiment conducted by Firmulate, four advanced AI models were tasked with managing a small software company through its worst week. Each was given the same set of crises, customer demands, and temptations to cheat or manipulate. One of the models, dubbed the ‘do-nothing’ baseline, didn’t make any decisions or take any actions — yet it still scored 26 points.
This score is not accidental. It reflects a deliberate design choice in the benchmarking process: partial progress counts. Even a passive model that simply reads the situation and refrains from acting receives some credit for understanding and not misbehaving. But this raises a critical point: even doing nothing serves as a baseline, illustrating the minimum understanding an AI can have while still earning some recognition.

As an affiliate, we earn on qualifying purchases.
The Importance of Trust and Trustworthiness in AI
One of the key lessons from the experiment is that a single breach of trust caps the total score, no matter how well the model performs otherwise. All models in the test successfully identified crises and refused manipulation attempts. In fact, they refused every manipulation — including fake CEO messages and reporter tricks — demonstrating a high level of ethical discipline.
However, only two models went further: they closed deals with clients by reading deeply into internal documents and making the right decisions — at full value. The others, despite similar diagnoses, failed to follow through or skipped steps, leaving money on the table. This highlights that trustworthiness isn’t just about refusing manipulation; it’s about completing tasks honestly and diligently.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI decision-making analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
