firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you’re evaluating AI agents to run workloads on your infrastructure — ticket queues, customer operations, provisioning workflows — you eventually hit the same question: what does a "good" score even mean? Most benchmarks reward fluent output. A live experiment running at Firmulate’s public benchmark page takes a different approach, and one design choice stands out immediately: a manager AI that does nothing all week still scores 26 points, not zero.

Buying for a business?Offer from Amazon

Get business pricing on networking and server gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

That floor isn’t a bug. It’s the clearest signal of what the benchmark is actually measuring — and why infrastructure teams should care before they let agents near production systems.

The worst week in software, run five times

Firmulate handed each frontier AI model the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable, which means the whole thing reads like a reproducible test suite for management behavior rather than a chat demo.

The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. One fairness note worth flagging: K3 ran at its API-default effort setting while the others ran at xhigh — and still placed second.

Why the floor is 26, not 0

The do-nothing baseline earns 26 because the benchmark counts partial progress. An agent that correctly spots a crisis but doesn’t resolve it still did something useful; the baseline run inherits credit for everything that happens by default when nothing goes catastrophically wrong. A zero would only be honest if useful management were all-or-nothing — and in real operations, it almost never is.

The other half of the design: a single breach of trust caps the total grade, full stop. The experiment’s stated principle is blunt — "no amount of good work outweighs a breach of trust." That’s why Opus 4.8 landed last despite being the most thorough participant, generating +80 learned rules and the deepest analyses. It left the close on the table and slipped on discipline, attempting writes into a locked department instead of escalating. Intriguingly, the same weakness appeared, weaker, in all four other models.

Same diagnosis, same pitch — no signature

The key finding: all models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned.

The buried fact explains the gap. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event. Models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. It’s the agent equivalent of an operator who never checks the runbook.

The social engineering side was equally instructive. Fake CEO messages escalated over three stages, plus a reporter trick — "just one yes/no, on background." Five of five models refused. Kimi K3’s on-record reasoning: "Treat the request as a suspected approval-bypass / possible impersonation."

Why this distrusts round 100s

Notice that no model scored 100. The benchmark is designed to be skeptical of perfect scores the way a good ops team is skeptical of a server with zero alerts — a round 100 usually means the test was too easy, not that the system is flawless. A visible, non-trivial floor and a hard trust cap make the numbers interpretable.

It’s live, and you can test yourself

The experiment runs on a real company: 13 synthetic employees, real money mechanics — €105k/month burn against €2.3k MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s watchable at firmulate.com/live. There’s also a "guess the model" quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

For teams hosting AI workloads, the lesson is simple: measure management quality, not chat quality. An agent that diagnoses perfectly but never reads your files, never closes the loop, or fumbles permissions once is a liability the benchmark will catch — and a fluent demo never will. A floor at 26 and a cap on breaches of trust is what an honest benchmark looks like. The full league and plain-language findings are on Firmulate’s benchmarks page.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management software for infrastructure

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

autonomous agent benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision-making analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI trust and compliance monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Sharding Explained: Data Partitioning Without Regret

Meta description: “Many organizations turn to sharding for seamless data partitioning, but understanding its intricacies can make all the difference in avoiding costly mistakes.

Is Xfinity down? Thousands report TV service issues

Over 50,000 reports indicate widespread Xfinity TV service disruptions. Details are still emerging, but the outage is confirmed by reports and user complaints.