
Fashion runs on judgment under pressure: a supplier slips, a customer walks, a rival makes a move, and someone has to decide what happens next. As AI takes on more business tasks, a polished chat demo cannot show how it will handle that kind of week. Firmulate’s experiment puts models inside a small company and watches what they do.
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A live company, with real stakes in the experiment
Firmulate is a public, watchable experiment in AI management. Its live company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. The workday is versioned, and the company has learned more than 680 playbook rules. Readers can follow the experiment at firmulate.com.
For its final Crucible League, in July 2026, five models faced the same small software company through its worst week: identical customers, crises and temptations. The published ranking was gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s principle is plain: partial progress counts, but a single breach of trust caps the total. No amount of good work outweighs a breach of trust.
Spotting trouble was not the same as finishing the job
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The finding captures a practical gap: “Same diagnosis, same pitch — no signature.” Recognizing the right move and carrying it through are different tests.
The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won at full price, worth €4,583 in monthly recurring revenue. In another test, fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.”
Opus 4.8 offers a revealing case. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. Kimi K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a relevant caveat when comparing the results.
From watching to testing your own playbooks
The league is a controlled public experiment, not a forecast of how any one company’s AI will behave. Its value is in making decisions visible: models can recognize a crisis, resist pressure and still fall short when a task requires follow-through or a sound escalation.
That is where the story turns from watching to acting. Enterprises can run a pilot against a read-only export of their own business, testing crisis scenarios against their own customers, pipeline and rules. The resulting board report includes model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each call.

For a fashion company considering AI across customer service, sales or planning, the useful question is not only whether a model can identify the right answer. It is whether it can act with discipline when the details are buried, pressure rises and the next step matters. Explore a Firmulate pilot for your business, and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
