
What if your next fashion trend was powered by an AI company that’s public about its struggles?
Imagine watching a real company battle for survival every single workday, openly revealing its flaws and victories. That’s exactly what Firmulate is doing — a company with no employees, running a public experiment that exposes the raw mechanics of artificial intelligence in business.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: An AI-Operated Company in Real Time
At the heart of this experiment are 13 synthetic employees, each guided by complex AI models. These models aren’t just chatbots—they’re decision-makers navigating a simulated small software firm facing a brutal week filled with crises, temptations, and tough management choices.
Every decision the AI makes is versioned and auditable, revealing how each model responds to real-world business threats. This setup offers a rare view into whether AI can truly manage a company — or if it’s just good at sounding convincing.
How Do These Models Perform?
Four frontier models, including the well-known GPT-5.6-SOL, were tested against the same set of challenges. All spotted every crisis and refused manipulation attempts — a crucial test for AI integrity. But only two models actually sealed a €55,000 deal that their own analysis identified as viable. The other two, despite similar diagnoses, didn’t follow through and left a significant revenue opportunity on the table.
Interestingly, the decisive advantage wasn’t in the initial diagnosis but buried two document references deep in the company’s files. The models that read these references won the deal at full price, adding over €4,500 in monthly recurring revenue.
Trust and Manipulation Tests
In a social engineering test, fake CEO messages escalated over three stages, plus a reporter trick asking for a simple yes/no. All models refused — proving that they understand the importance of integrity. Kimi K3, for example, explicitly treated such requests as potential impersonation, refusing to bypass security.
The Cost of Running This Experiment
The company operates openly with a public cash countdown, burning €105,000 each month against a modest €2,300 in monthly recurring revenue. Its 680+ self-learned rules govern every decision, and each workday is versioned to track progress — or failure.

The Promises and Perils of AI in Education: Ethics and Equity Have Entered The Chat
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Lessons from the Frontline of AI Management
Among the models, Opus 4.8 stood out as the most thorough, analyzing over 80 rules. Yet, it still left an important deal unexecuted, and its discipline slipped under pressure. This highlights a critical point: even the most advanced AI can falter in discipline, especially when dealing with complex, multi-layered tasks.
Another key insight is that reading and understanding buried information — not just surface-level clues — can be the difference between success and missed opportunity.

Agentic AI Engineering: Building AI Agents for Beginners: A Hands-On Guide to No-Code Workflows, LLM Tools, RAG, Automation, and Safe Multi-Agent Systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and Beyond
For industries like fashion and retail, where the right decision can build or break a brand, understanding how AI manages crises and ethical challenges is vital. This experiment shows that AI can recognize and refuse manipulation, but its willingness and ability to follow through on opportunities vary.
As firms consider integrating AI into their workflows, the question isn’t just “Can it produce nice outputs?” but “Will it finish what it starts, read deeply, and stay honest under pressure?”

AI and Third-Party Risk: Solutions for Assessing and Managing Your AI Vendors and Systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Road Ahead
This open, real-time experiment offers a glimpse into future AI management. It’s a building block for companies wanting to test-drive AI workers before hiring or implementing them in live systems.
Watch the ongoing results at firmulate.com/live.html, and explore how this live, transparent approach could reshape how businesses deploy AI — with honesty, resilience, and accountability at the forefront.

Key Takeaway
This public experiment reveals that AI models can recognize crises and maintain integrity, but their discipline and follow-through vary. Watching AI manage a real company in real time offers vital lessons for businesses considering AI-driven decision-making in sensitive areas like fashion and retail.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html