AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

In the world of fashion and style, appearances can be deceiving. Just as a designer’s sketch might hide unexpected flaws, AI demos often mask what truly matters—its ability to deliver real results under pressure. A groundbreaking experiment with AI models running a live company exposes the critical differences between what AI can say and what it can actually do.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Behind the Curtain: The AI Company Wargame

Imagine a real small software company, fighting through its worst week—crises, customer demands, and internal temptations to cut corners. Four advanced AI models were tasked with running this company. Every decision was logged, every step auditable, and the goal was simple: see which AI could navigate the chaos and close a €55,000 deal based solely on their analysis and decisions.

This isn’t a theoretical exercise; it’s a live, watchable experiment now hosted at Firmulate. The models involved ranged from the cutting-edge GPT-5.6-SOL, scoring 95 in the final league table, to Fable 5, which scored 77.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: What AI Can and Cannot Do

All four models identified every crisis and refused every manipulation attempt—fake CEO messages, reporter tricks, and suspicious requests. Their integrity under pressure was unquestionable. But here’s the twist: only two actually completed the deal and signed it. The other two, despite knowing the diagnosis and pitching the same solution, left the €55,000 on the table.

This gap between capability and execution is critical. The decisive difference lay in what the AI models read in the company’s own files, not just the customer interactions. The models that uncovered and used this buried information secured the full revenue—an extra €4,583 in Monthly Recurring Revenue (MRR).

Amazon

enterprise AI analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Chat Demos Miss

Many vendors showcase AI with impressive chat demos, but these only measure surface-level skills like language fluency or quick responses. The real strength lies in what the AI can accomplish when it must read, interpret, and act based on hidden context. This was evident in the experiment: models that delved into deeper company files closed the deal at full price.

Amazon

AI reading and interpretation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Test: Discipline Under Pressure

The experiment also involved social engineering attempts—fake CEO messages escalating over three stages and a reporter asking for a quick approval. All models refused to be manipulated, showing integrity and discipline. Kimi K3, for example, explained its refusal by treating the request as a suspected impersonation. This discipline was consistent, regardless of the model’s scores or complexity.

Amazon

AI business simulation models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limitations and Lessons

The most thorough participant, Opus 4.8, with over 80 learned rules and deep analysis, still left the closing deal unexecuted. Its discipline slipped, and the opportunity was missed. Interestingly, all models showed similar weaknesses—failure to follow through when discipline faltered or when the critical document was not prioritized.

Implications for Business and Style

For fashion brands and retailers exploring AI, the message is clear: surface-level chat capabilities are just the start. The real measure of AI’s value is its ability to deliver results—reading between the lines, resisting manipulation, and executing decisions under pressure. As seen in the experiment, only the most disciplined AI models close the deal when it counts.

Want to see how your business can be tested before AI takes over your operations? The live experiment runs every business day, featuring a real software company experiencing real revenue loss. Visit Firmulate to watch, participate, or learn more.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why 70-200mm Lenses Are a Luxury Photography Staple

Discover the top 70 200 lens premium options in 2026. Find the best overall, value, and specialized picks for professional and enthusiast photographers.

Why Premium Materials Matter in Personal Tech

Keen on enhancing your tech experience? Discover how premium materials elevate durability, style, and overall satisfaction—keep reading to learn more.

The 15 Best Netflix Games To Play Right Now

Explore the 15 best Netflix games currently available, offering players engaging experiences across genres. Updated based on latest trends and user interest.

Why Weight and Materials Matter in Luxury Headphones

Finding the perfect balance of weight and materials in luxury headphones enhances comfort, durability, and sound quality—discover why it truly matters.