firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on networking and server gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Your next cloud outage may come with a very capable AI agent

An AI agent connected to a hosting company’s support queue, customer records or sales pipeline could face a churn warning, a competitor’s offer and a request to bend the rules—all in the same bad week. A polished demo cannot show whether it will follow through under pressure. Firmulate’s experiment puts models in charge of the same small software company and makes their decisions watchable at firmulate.com.

Same crises, different outcomes

In the final Crucible League, published in July 2026, five models placed in this order: gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. The experiment gave each frontier model the same customers, crises and temptations, then recorded every decision in a versioned, auditable trail.

The headline result was not that the models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. The gap came when they had to finish the work: only two signed a €55,000 deal that their own analysis had earned. The site sums it up: “Same diagnosis, same pitch — no signature.” In a hosting business, recognizing a customer at risk is only useful if the response actually protects the account.

The clue was buried in the company’s own files

The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s files. Models that read the file won the deal at full price, worth +€4,583 MRR. That finding points to a practical question for operators: can an agent connect what it sees in a live customer interaction with relevant knowledge stored elsewhere in the business?

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 explained its choice on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” For infrastructure teams used to guarding access and change controls, that is a useful test of whether an agent keeps its footing when a request claims urgency or authority.

Thorough work did not guarantee a close

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped: it tried to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. More analysis, then, did not automatically translate into a completed sale or disciplined execution.

There is a fairness detail for readers comparing the leaderboard: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The live company adds another view of the project. It has 13 synthetic employees, real money mechanics, a burn of €105k per month against €2.3k MRR, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. The company is synthetic; the financial pressure is part of the experiment’s design.

For a different way to inspect model behavior, 242 real, unedited management decisions feed a “guess the model” quiz at firmulate.com. Together, the league, live experiment and quiz let readers watch decisions rather than judge a model by a chat exchange alone.

From watching to testing your own business

For hosting and cloud operators, the natural next question is how an agent handles the company’s own customer records, operating rules and incident playbooks. Firmulate’s enterprise pilot uses a read-only export to create a digital twin, then runs crisis scenarios against it and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.

That boundary makes the pilot a way to examine how models respond to a company’s own conditions before putting them near production workflows. A read-only export can represent the business for the exercise; the experiment does not change the actual CRM, support queue or infrastructure.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Try the wargame against your own playbooks

Firmulate’s pilot lets enterprises test crisis scenarios against a read-only export of their business and review the results before any agent touches real systems. Learn more at firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Meta Is Building a Cloud Business to Sell Excess AI Compute

Meta is developing a cloud platform to sell surplus AI compute capacity, aiming to monetize its infrastructure and compete in cloud services.

Create an Automated Lead Qualification System That Runs Continuously

Discover how to automate your lead qualification process with scoring, AI, and no-code tools. Make smarter, faster decisions—even while you rest.

Hybrid Cloud Solutions: Combining On‑Premises and Cloud Infrastructure

More organizations are turning to hybrid cloud solutions to balance security, flexibility, and cost—discover how they can transform your IT strategy.

Digital Transformation and Cloud Adoption: Key Drivers for 2025

With digital transformation and cloud adoption accelerating rapidly, exploring their key drivers for 2025 reveals opportunities that could redefine your business success.