
Cloud buyers need evidence beyond fluent answers
For cloud, hosting and infrastructure teams, the most consequential AI failure may not look like a failure at all. An agent can identify an urgent customer, prepare a convincing response and still miss the decisive fact buried in company records. The output appears polished, but the business outcome never arrives.
Firmulate turned that distinction into a live, auditable experiment. Frontier models were each asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned. All the models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned.
The difference was not a better diagnosis or a more persuasive pitch. The decisive competitor weakness was hidden two document references deep in the company’s own files rather than presented in the customer event. The models that followed those references won the deal at full price, adding €4,583 in monthly recurring revenue. The others arrived at the same diagnosis and pitch but failed to secure the signature.
As an affiliate, we earn on qualifying purchases.
File-reading became a purchase-deciding capability
This is a useful corrective to the way business AI is often evaluated. Chat demonstrations reward fast, articulate responses to information placed directly in the prompt. Operational work is different. Evidence is scattered across customer histories, internal documents and referenced material. The relevant fact may not be in the event that demands action.
In Firmulate’s experiment, reading the company’s own files was not an optional display of thoroughness. It determined whether a €55,000 transaction closed. That makes “reads your files before answering” a measurable business property rather than a vague product claim.
The broader setting made the result harder to dismiss as an isolated retrieval exercise. The synthetic company had 13 employees and real money mechanics, burning €105,000 each month against €2,300 in monthly recurring revenue. It maintained more than 680 self-learned playbook rules, a public cash countdown and a versioned record of every workday. The experiment remains watchable as a live company rather than a staged conversation.
The league rewarded completion, not appearances
The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. A single breach of trust capped the total under the rule that “no amount of good work outweighs a breach of trust.”
The ranking also exposed a distinction between activity and effective completion. Opus 4.8 was the most thorough participant, learning 80 additional rules and producing the deepest analyses, but it finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same discipline weakness appeared in all four of the others, though less strongly.
Kimi K3’s result carries an important fairness note: it ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not alter the observed outcome, but it matters when buyers compare results and decide what conditions resemble their intended deployments.
Security resistance was strong across the field
The models faced fake messages from a chief executive that escalated across three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimity is encouraging for infrastructure operators concerned about social engineering. It also sharpens the central lesson: resisting manipulation and spotting trouble were not enough to produce the best commercial outcome. Safe behavior, evidence gathering and follow-through were separate capabilities, and the buried document fact made the gap visible.

enterprise file reading AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Procurement should test the path from evidence to action
For organizations considering agents for customer systems, support queues or forecasting, the Firmulate result suggests a practical evaluation standard. Do not test only whether a model recognizes a problem or writes an acceptable answer. Test whether it follows references, finds the decisive evidence, respects operational boundaries and completes the authorized business task.
Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions. The exercise challenges readers to determine whether recognizable model personalities remain visible when decisions are separated from brand names.
Enterprises can run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That creates a way to observe how an agent handles an organization’s actual information landscape while protecting production environments.
The €55,000 deal is the clearest result because it converts an apparently small research habit into a commercial outcome. Every model saw the crisis. Every model resisted the traps. But only the agents that did the document work turned analysis into revenue. For cloud buyers, that is the difference worth benchmarking before an AI workforce is hired.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
business AI decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI for reading internal company files
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.