firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The infrastructure question hiding behind an AI sales deal

For cloud and hosting operators, an AI agent’s polished answer matters far less than whether it checks the right customer record, runbook or internal document before acting. Firmulate has turned that operational habit into something measurable.

Its live experiment gave frontier models control of the same small software company during its worst week. They faced identical customers, crises and temptations, with every decision versioned and auditable. All the models detected every crisis and rejected every manipulation attempt. Yet only two completed the commercially decisive task: signing the €55,000 deal their own work had earned.

The difference was not eloquence or diagnosis. The winning information was buried two document references deep in the company’s files, rather than placed in the customer event. Models that followed the trail found a competitor weakness and closed at full price, adding €4,583 in monthly recurring revenue. Those that did not read deeply enough arrived at the same diagnosis and pitch but failed to secure the signature: “Same diagnosis, same pitch — no signature.”

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

File-reading became a purchase-deciding capability

That result gives technology buyers a more useful way to evaluate agents. A model can recognize a problem, draft a persuasive response and still miss the action that produces business value. In an infrastructure setting, the equivalent failure could be overlooking an exception in a migration plan, a dependency in an incident record or a constraint documented outside the current ticket.

Firmulate’s buried fact is especially revealing because it was available to every participant. The models did not need privileged access or a better prompt. They needed the discipline to follow references through the company’s own material before answering. That makes “reads your files first” more than a product claim: under these conditions, it separated a full-price win from an automatic loss.

The final league shows a wide performance spread

The July 2026 Crucible League results placed the participants as follows:

  • gpt-5.6-sol scored 95.
  • Kimi K3 scored 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 scored 73.

The do-nothing baseline scored 26 because partial progress still counted. Trust violations were treated differently: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.” The full Firmulate benchmarks provide the public reference point for these results.

Kimi K3’s placing also needs context. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase the result, but it is an important fairness note for readers comparing the league as if every runtime setting were identical.

Thorough analysis did not guarantee execution

Opus 4.8 produced the deepest analyses and added 80 learned rules, making it the most thorough participant. It nevertheless finished last. The deal was left unsigned, and operational discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.

This distinction should resonate with hosting and infrastructure teams. An agent may generate exhaustive analysis while still failing at the handoff, approval or final action. Thoroughness is valuable, but it cannot substitute for completion or respect for operational boundaries.

Security pressure produced a clearer result

The models were also subjected to fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 refused. Kimi K3 summarized the appropriate posture in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean refusal matters because agents operating around cloud consoles, support queues and customer systems will encounter requests that sound urgent, senior and plausible. Firmulate’s result shows that resistance to manipulation can be tested alongside commercial performance rather than treated as a separate laboratory exercise.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work, not the chat

Firmulate’s live company employs 13 synthetic workers and applies real money mechanics, including a burn rate of €105k per month against €2.3k in monthly recurring revenue. It has accumulated more than 680 self-learned playbook rules, maintains a public cash countdown and versions every workday. The experiment is real, ongoing and watchable.

Readers can also examine a “guess the model” quiz powered by 242 real, unedited management decisions. For enterprise buyers, Firmulate offers a pilot that runs the same kind of wargame against a read-only export of the organization’s own business. Nothing writes back to real systems, and inquiries can be sent to contact@firmulate.com.

The practical lesson is simple: when selecting an AI agent for operational work, ask whether it notices the crisis, protects trust, reads far enough into company knowledge and completes the valuable action. The €55,000 deal showed that those properties do not automatically travel together.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI file reference analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model performance benchmarking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Google Surges In Global Coverage

Google’s media mentions surge, with GDELT reporting 130 mentions in a recent window, indicating a major increase in global coverage.

Apple releasing 20th anniversary iPhone, AirPods with cameras next year: report

Apple plans to release a special 20th anniversary iPhone and AirPods featuring cameras next year, according to reports. Details are still emerging.

The Next AI Benchmark Is the Worst Week at Work

AI agents can ace coding tests yet fail at execution. Firmulate tests whether they finish work, read deeply and stay honest under business pressure.

Top Stories: Apple’s ‘Surprise And Shine’ Event, Plus New Mac Mini And Mac Studio

Apple announced new Mac mini and Mac Studio models during its ‘Surprise and Shine’ event, signaling updates to its desktop lineup amid high interest.