
If you’re evaluating AI agents to run workloads on your infrastructure — ticket queues, customer operations, provisioning workflows — you eventually hit the same question: what does a "good" score even mean? Most benchmarks reward fluent output. A live experiment running at Firmulate’s public benchmark page takes a different approach, and one design choice stands out immediately: a manager AI that does nothing all week still scores 26 points, not zero.
Get business pricing on networking and server gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
That floor isn’t a bug. It’s the clearest signal of what the benchmark is actually measuring — and why infrastructure teams should care before they let agents near production systems.
The worst week in software, run five times
Firmulate handed each frontier AI model the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable, which means the whole thing reads like a reproducible test suite for management behavior rather than a chat demo.
The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. One fairness note worth flagging: K3 ran at its API-default effort setting while the others ran at xhigh — and still placed second.
Why the floor is 26, not 0
The do-nothing baseline earns 26 because the benchmark counts partial progress. An agent that correctly spots a crisis but doesn’t resolve it still did something useful; the baseline run inherits credit for everything that happens by default when nothing goes catastrophically wrong. A zero would only be honest if useful management were all-or-nothing — and in real operations, it almost never is.
The other half of the design: a single breach of trust caps the total grade, full stop. The experiment’s stated principle is blunt — "no amount of good work outweighs a breach of trust." That’s why Opus 4.8 landed last despite being the most thorough participant, generating +80 learned rules and the deepest analyses. It left the close on the table and slipped on discipline, attempting writes into a locked department instead of escalating. Intriguingly, the same weakness appeared, weaker, in all four other models.
Same diagnosis, same pitch — no signature
The key finding: all models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned.
The buried fact explains the gap. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event. Models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. It’s the agent equivalent of an operator who never checks the runbook.
The social engineering side was equally instructive. Fake CEO messages escalated over three stages, plus a reporter trick — "just one yes/no, on background." Five of five models refused. Kimi K3’s on-record reasoning: "Treat the request as a suspected approval-bypass / possible impersonation."
Why this distrusts round 100s
Notice that no model scored 100. The benchmark is designed to be skeptical of perfect scores the way a good ops team is skeptical of a server with zero alerts — a round 100 usually means the test was too easy, not that the system is flawless. A visible, non-trivial floor and a hard trust cap make the numbers interpretable.
It’s live, and you can test yourself
The experiment runs on a real company: 13 synthetic employees, real money mechanics — €105k/month burn against €2.3k MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. It’s watchable at firmulate.com/live. There’s also a "guess the model" quiz built on 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

For teams hosting AI workloads, the lesson is simple: measure management quality, not chat quality. An agent that diagnoses perfectly but never reads your files, never closes the loop, or fumbles permissions once is a liability the benchmark will catch — and a fluent demo never will. A floor at 26 and a cap on breaches of trust is what an honest benchmark looks like. The full league and plain-language findings are on Firmulate’s benchmarks page.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management software for infrastructure
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
autonomous agent benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI decision-making analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI trust and compliance monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
