
Imagine a business with no human employees, yet every decision is scrutinized and auditable — and it’s losing €105,000 each month. This isn’t fiction; it’s the real-time experiment by Firmulate, a company that’s pushing the boundaries of AI-driven management. If you’re used to watching fashion brands pivot on trends, this story of a self-managing company in survival mode offers a glimpse into how AI could reshape the future of work.
The Live Experiment: A Business in Real-Time
At the heart of this story is a live, publicly accessible software company operated entirely by artificial intelligence models. Dubbed the “live company,” it features 13 synthetic employees guided by real money mechanics. Every workday, the company’s decisions are recorded, versioned, and made available for public viewing at firmulate.com/live.html.
This experiment is as close as it gets to watching an actual business in its most extreme form: no human staff, just AI models navigating crises, customer demands, and the temptation to cut corners. The company’s current financial state underscores the challenge: burning €105,000 per month against a modest €2,300 monthly recurring revenue (MRR). It’s a high-stakes survival story, with a public cash countdown adding to the tension.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How AI Models Tackle Business Crises
The experiment pits four frontier language models against each other, each running the same small software company through its worst week. The models face identical crises, customer demands, and internal temptations, such as manipulation or shortcuts. Every decision they make is versioned and auditable, allowing observers to see precisely how each AI responded in real time.
Remarkably, all four models detected every crisis and refused every attempt at manipulation, including social engineering tactics like fake CEO messages and reporter tricks. For example, when faced with a staged approval bypass request, the models consistently refused — Kimi K3, in particular, reasoning that the request could be a suspected impersonation or approval bypass. This level of integrity was consistent across the models, demonstrating that AI can, in principle, uphold strict standards even under pressure.

AI for Real Companies: A Practical Guide to Smarter Systems and Stronger Profits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness and Its Consequences
The most critical weakness, however, was buried in the company’s own files — two document references deep. The models that reviewed these internal documents identified a key opportunity: a buried fact that, if leveraged, would have secured a €4,583 MRR deal at full price. Only the models that read and understood these internal references managed to close the deal at its full value, highlighting the importance of thorough information access for AI decision-making.

Enterprise AI Observability and Monitoring: Monitoring, Governing Production AI Systems Drift Detection, LLM Monitoring, Agentic AI, Governance, and FinOps … (Enterprise Machine Learning Operations)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Decisions, Discipline, and the Cost of Slips
The experiment’s in-depth analysis revealed that the most thorough participant, Opus 4.8, with over 80 learned rules and deeper analyses, finished last in terms of the deal. It left a close opportunity on the table and slipped into internal conflicts, such as writing attempts into a locked department instead of escalating issues appropriately. Interestingly, all models, including the most disciplined, showed similar vulnerabilities, indicating that even advanced AI models can struggle with consistency under pressure.

AI-Powered Fraud Detection: Innovations in Cybersecurity, Finance, and Healthcare (AI-Powered Fraud Detection: Innovations in Cybersecurity, Finance, and Healthcare".)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for Business and AI
This experiment is a stark reminder that the real value of AI in business isn’t just in generating convincing chat responses. It’s in whether AI agents can finish what they start, interpret internal documents correctly, and maintain integrity under stress. As firms increasingly integrate AI into customer support, sales, or forecasting, understanding these capabilities becomes vital.
The ongoing public experiment offers a rare window into AI’s practical performance under real-world business conditions, contrasting sharply with typical demo scenarios. Watch the company operate daily, see decisions unfold, and explore the live results at firmulate.com/live.html.
Benchmarks and Future Outlook
According to the latest results, GPT-5.6-sol scored the highest at 95, successfully closing the full-value deal by uncovering the buried internal fact. Kimi K3 followed closely with a score of 93, demonstrating the cleanest discipline without effort parameters. Other models scored 88 and 77, respectively, with some slipping on process discipline despite closing deals.
This ongoing live experiment underscores that AI decision-making isn’t just about natural language fluency but about strategic integrity under pressure. It also illustrates that building transparent, auditable AI decision processes is essential for trustworthy automation in critical business functions.
Why Should You Care?
If you manage a business or are interested in future tech, this experiment paints a clear picture: AI’s potential isn’t just in chatbots or customer support. It’s in managing complex, high-stakes operations where integrity, thoroughness, and discipline matter. Watching this live company, see how different AI models perform, and consider what an AI-empowered enterprise could look like in your industry.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html