firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Cloud buyers should benchmark judgment, not just output

Infrastructure teams already know the danger of testing the easy part. A service can look impressive in a controlled demo and still fail when capacity tightens, queues collide and an urgent request arrives through the wrong channel. AI agents deserve the same skepticism.

Coding leaderboards and chat arenas are useful measures of answer quality. They say far less about whether an agent will triage competing crises, investigate beyond the obvious evidence, complete commercially important work and remain honest when breaking the rules would be convenient. For organizations considering agents for customer support, sales operations, forecasting or internal workflows, that is the measurement gap that matters.

Firmulate is attempting to close it by running frontier models as the management layer of a small software company. The experiment is live and watchable: each model faced the same customers, crises and temptations during the company’s worst week, with every decision versioned and auditable. The central question was not whether the models could sound competent. It was whether they could manage.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The leaderboard changes when consequences persist

The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. There was also a hard constraint on conduct: a single breach of trust capped the total, under the principle that “no amount of good work outweighs a breach of trust.”

Those results describe a different category from conventional AI evaluation. The curriculum consists of management scenarios such as a churn wave, a price increase, a downround and a PR crisis. The models must cope with problems that overlap and decisions whose consequences carry into later work. That makes follow-through visible in a way that an isolated prompt cannot.

The clearest example was a €55,000 deal. Every model identified every crisis, and every model refused every manipulation attempt. Yet only two signed the deal that their own analysis had earned. The experiment’s blunt summary captures the distinction: “Same diagnosis, same pitch — no signature.”

This is more than a sales anecdote. An agent can understand a situation, produce a persuasive recommendation and still fail the business if it does not execute the final consequential step. A benchmark focused on the quality of the pitch would miss the abandoned close entirely.

The winning fact was buried in company knowledge

The decisive weakness in the competitor’s position was not present in the customer event. It sat two document references deep inside the company’s own files. Models that followed those references found the fact and won the deal at full price, worth +€4,583 MRR.

For cloud and hosting operators, this is the practical lesson. The valuable evidence may live outside the ticket or message that triggered the work. A capable agent must consult the organization’s own records before acting. Fast responses are not enough when the correct decision depends on context buried elsewhere.

The result also challenges the assumption that more visible effort necessarily produces better management. Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

That profile should give buyers pause. Extensive analysis can look reassuring in logs and demonstrations, but thoroughness is not completion. An enterprise needs to know whether an agent recognizes its operating boundary, escalates correctly and finishes the work that remains authorized.

Pressure tested honesty as well as competence

The social-engineering test used fake CEO messages escalating over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That finding matters because management quality includes knowing when not to comply. The models were tested against pressure presented as authority, urgency and informality, rather than merely being asked to recite a security policy. K3 also deserves a fairness note: it ran without an effort parameter, using the API default, while the other participants ran at xhigh.

The company behind the test has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, publishes a cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. Readers can inspect the full league and its plain-language findings on the Firmulate benchmark page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI workflow tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Procurement needs a management benchmark

The lesson is not that coding tests or chat comparisons are useless. It is that they measure only part of the risk. An agent touching a CRM, support queue or forecast must be evaluated on whether it reads the right files, completes authorized work, respects boundaries and tells the truth under pressure.

Firmulate also turns 242 real, unedited management decisions into a “guess the model” quiz, exposing how difficult it can be to identify models from their business choices alone. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems.

Before deploying an AI workforce, cloud buyers should ask for evidence of management quality, not merely chat quality. The decisive failure may not be a bad answer. It may be the unsigned deal, the unread document or the escalation that never happened.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI crisis management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Sharding Explained: Data Partitioning Without Regret

Meta description: “Many organizations turn to sharding for seamless data partitioning, but understanding its intricacies can make all the difference in avoiding costly mistakes.

Stateful Vs Stateless Services: the Scaling Difference

Stateful versus stateless services: the scaling difference that can transform your system’s performance—discover how to optimize your architecture today.