
Get business pricing on networking and server gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Your next infrastructure decision may be which AI gets access
AI agents are moving toward work that touches customer records, support queues and forecasts. For cloud and hosting teams, a polished demo is a thin basis for trust. What matters is whether an agent can read the evidence, make a sound decision and follow through when a real business is under pressure. Firmulate has put that question to a public test: a live software company run by AI models, with crises, customers and money mechanics.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company, not a chat prompt
In Firmulate’s Crucible experiment, frontier models were given the same small software company to run through its worst week. They faced the same customers, crises and temptations, and every decision was versioned and auditable. The company has 13 synthetic employees, burns €105,000 a month against €2,300 in monthly recurring revenue, and publishes its cash countdown. It runs every business day; visitors can watch the experiment.
The final July 2026 league table put gpt-5.6-sol first at 95, followed by Moonshot’s Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26. Firmulate’s rule is stark: partial progress counts, but one breach of trust caps the total; “no amount of good work outweighs a breach of trust.”
AI document access and management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The difference between spotting a problem and finishing the job
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. That gap—“Same diagnosis, same pitch — no signature”—is the practical test for anyone considering agents for operational work: recognizing the right answer does not guarantee that the system will complete the task.
The deal hinged on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The result makes document access and follow-through part of the management test, not just the quality of an agent’s written analysis.
Kimi K3 finished second overall, ahead of three of the four Western frontier models in the field. It found the buried security needle, won the deal, saved the churning customer and resisted all three baits, with one deviation—the cleanest discipline in the field. In a social-engineering sequence, fake CEO messages escalated through three stages, followed by a reporter’s “just one yes/no, on background” request. All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a different lesson. It was the most thorough participant, with 80 learned rules and the deepest analyses, but ranked last. It left the deal unsigned and attempted writes in a locked department instead of escalating. A weaker version of that discipline problem appeared in all four. More analysis, by itself, did not ensure a better outcome.
AI security and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test before you trust
Firmulate says the workday decisions are drawn from a company with 680+ self-learned playbook rules, and offers a quiz built from 242 real, unedited management decisions. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems. The full benchmark and plain-language findings are public.
Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

AI crisis management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Run your own trial
The league is open: Kimi K3 came close to the leader and beat three of four Western frontier models, while the strongest analyses did not always turn into completed work. For infrastructure and hosting teams weighing AI agents, choosing from a leaderboard or a chat demo alone is a bet. Test the model against your own operational pressures before entrusting it with the work.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
