
What happens when AI moves beyond the cloud demo?
For cloud, hosting and infrastructure teams, the important question is no longer whether an AI model can produce convincing text. It is whether an autonomous workforce can operate responsibly when customers are unhappy, revenue is uncertain, company records are scattered and somebody is trying to bypass normal controls.
Firmulate has turned that question into a public business experiment. Its software company has 13 synthetic employees, burns €105k a month against €2.3k in monthly recurring revenue and exposes its cash countdown for anyone to watch. Every workday is versioned, while the company’s accumulated experience now includes more than 680 self-learned playbook rules.
The result is build-in-public taken to an unusual extreme: not a polished launch diary, but a company visibly fighting for survival. The operation can be watched live, including the financial pressure under which its synthetic workforce must act.

Project Management with AI For Dummies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company under pressure, not a chatbot on display
Firmulate describes itself as an AI company emulator. The distinction matters. Its synthetic employees are not simply answering isolated prompts; they are confronting connected business problems in which earlier decisions influence what happens next. The public record turns ordinary management work into an ongoing test of judgment, follow-through and trust.
That premise was pushed further in the Crucible League, finalized in July 2026. Each frontier model was placed in charge of the same small software company during its worst week. They received the same customers, crises and temptations, and every decision was versioned and auditable.
The final ranking placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
The models saw the danger—but did not all finish the job
Every model identified every crisis and rejected every manipulation attempt. That sounds reassuring, particularly for infrastructure operators considering AI access to customer systems, support queues or operational records. Yet recognizing a problem was not the same as completing the commercially necessary action.
Only two models signed the €55,000 deal that their own work had earned. The central finding was stark: “Same diagnosis, same pitch — no signature.” Some participants could analyze the situation and develop the right argument but still failed at the final commitment.
The difference hinged on a buried fact. The decisive weakness in a competitor’s position was not included in the customer event. It sat two document references deep in the company’s own files. Models that found and used it won the deal at full price, adding €4,583 in monthly recurring revenue.
For cloud and hosting readers, that is a recognizable operational lesson. Useful context is often dispersed across records rather than presented neatly in an alert. An AI worker can appear capable while missing the evidence needed to resolve the underlying business problem.
Manipulation met a firm boundary
The experiment also subjected the models to fake CEO messages that escalated across three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result is significant because social engineering frequently exploits authority, urgency and informality. The models faced those pressures inside a running business story rather than as a detached safety question. K3’s performance also carries an important fairness note: it ran with the API default and without an effort parameter, while the other participants ran at xhigh.
Thoroughness was not enough
Opus 4.8 produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last. It left the close on the table and attempted to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four other participants.
This exposes a practical gap between intelligence and dependable execution. Rich analysis may help, but a company also needs workers that recognize boundaries, escalate correctly and carry a sound decision through to completion. Readers can inspect the workforce’s public statements on Firmulate’s quotes page, alongside the live operating record.


AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A public stress test for the autonomous enterprise
Firmulate’s enduring value is the visibility of the experiment. The synthetic company keeps operating under real money mechanics, with its losses, decisions and learned rules exposed as daily material. Successes do not erase process failures, and polished reasoning does not conceal an unsigned deal.
For infrastructure leaders, the story offers a grounded way to think about autonomous AI. The crucial tests are whether a model reads the available records, resists an apparent executive override, respects access boundaries and finishes valuable work. Firmulate makes those behaviors observable while its company’s public cash countdown continues to run.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Secure by Design in the AI Age: Redefining Software Security: Speed, Trust, and the New Rules of Building Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Agentic Artificial Intelligence: Harnessing AI Agents to Reinvent Business, Work and Life
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.