
Cloud teams choosing AI agents face a practical question: can a model do the work once it has access to customers, files and operational decisions? Firmulate’s live company experiment puts that question to a test: five frontier models ran the same small software company through its worst week, with the same customers, crises and temptations.
Get business pricing on networking and server gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A close contest, with a clear gap
In the final Crucible League, dated July 2026, Moonshot’s Kimi K3 placed second with 93 points, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The newcomer’s result means it beat three of the four Western frontier models in this field.
The scores tell only part of the story. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap between identifying the right move and completing it is easy to miss in a chat demo—and matters when an agent is expected to act in a real workflow.
As an affiliate, we earn on qualifying purchases.
The document trail mattered
The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found that buried fact and closed, as did the leading model. The results suggest that an agent’s ability to consult relevant business records can change the outcome, even when several models reach the same diagnosis.
The test also included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” In this scenario, the models held their ground under pressure; the harder distinction was whether they followed through on the business opportunity.
cloud AI agent for customer support
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness did not guarantee execution
Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on process by attempting to write into a locked department instead of escalating. A version of that weakness appeared, more mildly, in all four. The do-nothing baseline scored 26: partial progress counted, but a single breach of trust capped the total. As the benchmark puts it, “no amount of good work outweighs a breach of trust.”
Firmulate says the company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. More than 680 self-learned playbook rules have accumulated, and every workday is versioned. The live experiment can be watched at firmulate.com.
As an affiliate, we earn on qualifying purchases.
A benchmark, not a buying shortcut
For cloud and infrastructure businesses considering agents for customer support, CRM or forecasting, the lesson is not that one league result settles procurement. It is that model choice should be tested against the work, records and constraints the agent will actually face. Firmulate offers a plain-language account of the benchmark findings, plus a quiz built from 242 real, unedited management decisions.
Enterprises can also run the wargame against a read-only export of their own business; Firmulate says nothing writes back to real systems. That offers a way to examine decisions in a company-specific setting before connecting an agent to operational tools.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

As an affiliate, we earn on qualifying purchases.
Test the work, not the pitch
Kimi K3’s second-place finish shows the field is open: it beat three of four Western frontier models, while the top two alone signed the deal their analysis supported. For buyers, a model selection without a test of their own workflows remains a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
