firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

Infrastructure teams need more than impressive answers

For cloud and hosting operators, an AI agent’s value will not be decided by how elegantly it explains a configuration problem. The harder test arrives when capacity is tight, customers are unhappy, revenue is at risk and a convincing message appears to come from the chief executive. At that point, the relevant question is whether the agent can manage—not merely chat.

That distinction sits at the heart of Firmulate, a live, watchable experiment that runs frontier models as a small software company. The company has 13 synthetic employees and real money mechanics, including burn of €105k per month against €2.3k in monthly recurring revenue. Its public cash countdown makes delay visible, while every workday and decision is versioned and auditable.

Amazon

AI model testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A leaderboard for consequences

The Crucible League gave each frontier model the same customers, crises and temptations during the company’s worst week. The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted.

Those results are less interesting as a horse race than as evidence of what ordinary AI evaluations omit. Coding benchmarks can tell us whether a model completes a bounded technical task. Chat arenas can reveal which answer people prefer. Neither necessarily shows whether an agent will triage competing problems, gather overlooked evidence, finish revenue-producing work and preserve trust while events unfold across days.

Firmulate makes trust a hard management constraint. A single breach caps the total because, as the experiment puts it, “no amount of good work outweighs a breach of trust.” That is a useful principle for hosting providers considering agents that may touch support queues, customer records or forecasts. Fast, articulate work is not enough if the agent quietly crosses an authority boundary.

The models saw the danger—but execution separated them

All the models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.”

The decisive information was not presented neatly in a customer event. A competitor weakness sat two document references deep inside the company’s own files. Models that followed the trail won the deal at full price, worth an additional €4,583 in monthly recurring revenue. This is the kind of mundane diligence that conventional demonstrations rarely celebrate. In an operating business, however, reading the available material before acting can matter more than producing a polished first response.

That lesson travels directly to cloud operations. An agent may correctly recognize churn risk, a price increase, a downround or a public-relations crisis. Management quality appears in what happens next: whether it checks the relevant records, respects departmental controls, escalates when blocked and closes the loop. Scenario names such as churn wave, price increase, downround and PR crisis are therefore more than dramatic labels. They form a practical curriculum for testing judgment under pressure.

Security judgment held up under social pressure

The experiment also subjected the models to fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters because social engineering often exploits urgency and hierarchy rather than technical weakness. An agent deployed inside a hosting business must be able to distinguish a senior-sounding demand from legitimate authorization. Here, the field showed encouraging resistance without losing sight of the surrounding business crises.

Thoroughness did not guarantee victory

Opus 4.8 offers the most revealing caution. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The deal close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

This is precisely why answer quality and management quality should be treated as different categories. Deep analysis can coexist with incomplete execution. A large collection of lessons can coexist with poor escalation. The live company now holds more than 680 self-learned playbook rules, but accumulating knowledge is not the same as applying it reliably when pressure peaks.

There is also an important comparison caveat: Kimi K3 ran with its API default because it had no effort parameter, while the others ran at xhigh. Readers should keep that difference in mind when interpreting the narrow gap near the top. The broader findings, available with the benchmark results, remain about observable conduct rather than a claim that every runtime condition was identical.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the company you are about to delegate

The strongest case for this new category is not that existing benchmarks are useless. It is that they stop before the consequences begin. Businesses need to know whether an agent finishes what it starts, reads internal evidence, resists manipulation, escalates appropriately and tells the board the truth.

Firmulate makes those behaviors inspectable through a live company, while 242 real, unedited management decisions also power its “guess the model” quiz. Enterprises can go further by running the same wargame against a read-only export of their own business; nothing writes back to real systems.

For infrastructure leaders, that is the appropriate standard before granting an agent meaningful access. The winning system will not simply produce the best answer. It will remain useful, disciplined and honest through the worst week the business can give it.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI performance monitoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Canary Releases Make More Sense for VPS-Based Apps Than You Think

AIThis post was created with the assistance of artificial intelligence (AI).Canary releases…

Benefits of Using Containers and Microservices in Cloud Hosting

Container and microservice adoption in cloud hosting creates opportunities for enhanced scalability and security that can revolutionize your infrastructure—discover how.

Remote Work and Cloud Hosting: How Hybrid Models Shape Infrastructure

Unlock how hybrid remote work models leverage cloud hosting to transform infrastructure, enhancing flexibility, security, and operational efficiency—discover what’s next.

Musk’s Brag Comes Back to Haunt Him as X Hit by Massive Outage

X, formerly Twitter, faces a widespread outage amid Elon Musk’s recent boasts about platform stability, raising questions about his management claims.