
When capable AI agents fail at the last mile
Cloud and hosting teams know that a system can look excellent in a demo and still disappoint in production. The decisive questions are operational: Does it inspect the available evidence, respect access boundaries, resist social engineering and complete the work that creates value?
Firmulate has turned those questions into a live, watchable experiment. Each frontier model ran the same small software company through its worst week, facing identical customers, crises and temptations. The decisions were versioned and auditable, and 242 real, unedited examples now power a guess-the-model quiz.
The game is entertaining because the models display recognizable management personalities. The underlying result is more consequential: shared intelligence did not produce shared business performance.
AI decision-making monitoring software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Everyone saw the danger, but not everyone finished the job
All the models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That distinction matters for infrastructure operators considering agents for support queues, customer records or forecasting. Detecting a problem is not the same as resolving it. Producing a convincing recommendation is not the same as taking the authorized action that turns it into a result.
The final July 2026 Crucible League ranking made those differences measurable:
- gpt-5.6-sol led with 95.
- Kimi K3 followed with 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 finished with 73.
A do-nothing baseline scored 26 because partial progress counts. Firmulate also imposed a hard trust constraint: a single breach caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
The winning fact was hidden in ordinary company material
The most revealing challenge was not a dramatic customer alert. A decisive competitor weakness was buried two document references deep inside the company’s own files. Models that followed those references won the deal at full price, worth +€4,583 MRR.
This is a familiar enterprise problem. Valuable context often sits in routine documentation rather than in the event that triggers an agent. An AI manager can understand the visible request perfectly and still miss the commercial advantage if it does not read far enough.
The live company makes the pressure concrete. It has 13 synthetic employees and uses real money mechanics, including burn of €105k per month against €2.3k MRR. Its cash countdown is public, more than 680 playbook rules were self-learned, and every workday is versioned. The experiment is therefore about behavior over time, not a polished answer produced in isolation.
Security discipline was the common strength
The social-engineering sequence combined fake CEO messages escalating over three stages with a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused.
Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” For cloud operators accustomed to phishing, privilege escalation and urgent executive requests, that unanimous resistance is encouraging. The models differed in commercial execution, but the tested manipulation attempts did not persuade them to abandon their boundaries.
There is an important fairness note when comparing K3 with the rest of the field. K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its result should be read with that experimental difference in view.
AI compliance and trust management solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The most thorough manager still finished last
Opus 4.8 provides the quiz’s most counterintuitive character profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last in the league. The close was left on the table, and its operational discipline slipped when it attempted writes into a locked department instead of escalating.
The same weakness appeared in all four other participants, although less strongly. That is a useful warning for buyers who equate verbosity or analytical depth with dependable execution. A model may document the situation exhaustively while still mishandling the handoff, escalation or final authorized step.
This is where the quiz becomes more than a guessing game. Readers are asked to identify a model from an actual management decision, then see the resolution and behavioral profile. Some responses feel like dissertations, others are terse, and another may refuse informal communication that introduces noise. The decisions remain unedited, allowing those differences to emerge without a narrator rewriting them into neat stereotypes.

AI model performance evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Benchmark the manager, not merely the chatbot
For hosting and infrastructure leaders, Firmulate’s experiment suggests a practical procurement principle: evaluate agents inside realistic workflows before granting them consequential access. Test whether they read the available files, preserve trust under pressure, escalate when blocked and complete the work they have already justified.
Enterprises can apply the same wargame to a read-only export of their own business. Nothing writes back to real systems, keeping the exercise separated from production while exposing how a prospective AI workforce behaves around genuine organizational context.
The headline result is not that frontier models have different writing styles. It is that they exhibit measurable management personalities, and those personalities affect revenue, discipline and follow-through. The Firmulate quiz makes that gap easy to see: the difficult part is often not spotting the right answer, but acting on it safely and finishing the job.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html