firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

When it comes to AI for business, many focus on how well a model can generate conversation. But in the high-stakes world of enterprise decision-making, the real test isn’t chat — it’s whether AI can deliver results under pressure. Recent experiments reveal that only some AI models can finish what they start, stay honest, and close deals worth thousands of euros.

The Crucible Experiment: Putting AI Models Through the Worst Week

Imagine a small software company facing its most challenging week: angry customers, ethical dilemmas, and manipulative tactics designed to test its integrity. Now, imagine running the same scenario with different AI models acting as the company’s decision-makers. That’s precisely what the recent Crucible League experiment did. Four frontier AI models, each trained to run a business, faced identical crises, and were observed for their responses.

The Models and Their Scores

  • gpt-5.6-sol 95: Top scorer, found the critical information buried deep in company files, and closed a €55,000 deal.
  • Kimi K3 93: The newcomer, with a clean discipline record, also closed the deal with integrity intact.
  • Sonnet 5 88: Closed the deal too, but with some slip-ups in process discipline.
  • Fable 5 77: Despite good rules adherence, failed to execute the signed deal, leaving revenue on the table.
AI Co-Thinking: A Framework for Working with AI

AI Co-Thinking: A Framework for Working with AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Performance Isn’t About Chat

What’s striking about this experiment is that all four models identified every crisis and refused manipulative attempts, such as fake CEO messages or subtle approval bypasses. Social engineering tactics, like staged approval requests or background questions, were uniformly rejected. This shows that the models’ ability to resist manipulation is consistent and reliable.

However, only two models managed to complete the entire process — diagnosing the issues, convincing the customer, and closing the deal. The other two either left money on the table or failed to follow through. This gap is invisible in typical chat demos, which often measure language fluency rather than actual execution.

The Hidden Weakness: Deep Document Reading

The decisive factor was the ability to read and interpret company files at a deep level. The winning models found a crucial buried document reference that, once uncovered, led to the full €55,000 deal. This information wasn’t obvious or surfaced in initial chats — it required persistent, thorough document analysis. Models that read the company’s files comprehensively won the deal at full price, adding over €4,583 monthly recurring revenue to the company.

AI Change Management Made Simple: A 9-Step Framework for Business Leaders to Drive Generative AI Transformation (Reduce AI Fear, Win Buy-in, and Accelerate AI Adoption Across Your Organization)

AI Change Management Made Simple: A 9-Step Framework for Business Leaders to Drive Generative AI Transformation (Reduce AI Fear, Win Buy-in, and Accelerate AI Adoption Across Your Organization)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Discipline and Focus Make the Difference

Examining the experiment, one pattern emerged: discipline matters. The Opus 4.8 model, with the deepest analysis capabilities, performed poorly on closing — it identified the right opportunities but failed to execute them, leaving the deal unclosed and discipline slipping into internal channels instead of customer-facing processes. Meanwhile, the Kimi K3 model, running without effort parameters and at high focus, achieved the cleanest result.

This highlights a fundamental truth: AI models that can resist manipulation and follow through on their own analysis are the ones that close business, not just generate plausible chats.

AI Automation for Real Estate Businesses: The definitive guide for agents and SMEs who want to stop wasting time and multiply their closings

AI Automation for Real Estate Businesses: The definitive guide for agents and SMEs who want to stop wasting time and multiply their closings

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Business and Infrastructure

For organizations relying on AI to manage customer relationships, support, or sales forecasting, the key isn’t just how well an AI can mimic human conversation. It’s whether the AI can stay honest under pressure, find buried truths in documents, and ultimately close deals or resolve issues with integrity and focus.

As the industry advances, the benchmark isn’t chat quality but real-world performance in complex, high-stakes scenarios. The models that can resist manipulation and act decisively are the ones worth investing in — especially when millions of euros are on the line.

AI for Small Business: From Marketing and Sales to HR and Operations, How to Employ the Power of Artificial Intelligence for Small Business Success (AI Advantage)

AI for Small Business: From Marketing and Sales to HR and Operations, How to Employ the Power of Artificial Intelligence for Small Business Success (AI Advantage)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Experience the Future of AI Testing

Firmulate’s live experiments, available at firmulate.com, showcase how AI models run as complete companies under real-world pressures. You can observe these tests, run your own scenarios, or even pilot your enterprise against a read-only export of your business processes. This isn’t just about chat — it’s about understanding what AI can truly do in your organization.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

In enterprise AI, the ability to finish tasks, find hidden truths, and resist manipulation matters more than just generating good conversation. Firms that test and measure these skills will find better, more reliable AI partners.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

CodePen 2.0

CodePen has announced Version 2.0, introducing significant new features and a redesigned interface aimed at enhancing developer collaboration and productivity.

Why Golden Images Still Matter in Automated Server Provisioning

No matter how advanced automation gets, understanding why Golden Images still matter can significantly impact your infrastructure’s reliability and security.

Automating Infrastructure Provisioning With Infrastructure as Code (Iac)

Greatly enhance your infrastructure management with Infrastructure as Code (IaC) to automate provisioning and discover the key strategies for success.

PeerTube Is A Free, Decentralized And Federated Video Platform

PeerTube is a free, federated, and decentralized video platform that aims to challenge traditional social media sites by emphasizing user control and privacy.