
Imagine running a restaurant where every decision, from ordering ingredients to handling customer complaints, is made by artificial intelligence. Now imagine watching that restaurant fight to stay open every single day, losing money while trying to meet impossible standards. That’s the reality of a pioneering experiment in AI management, available for anyone to observe live.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Live Company in Action: A New Kind of Business Experiment
At the forefront of AI experimentation, a real software company operates with a twist — it has no human employees and is managed entirely by artificial intelligence models. This company, watched daily at firmulate.com/live, is a high-stakes laboratory for understanding how AI can handle real-world business crises, ethical dilemmas, and decision-making under pressure.
How It Works
The company runs with 13 synthetic employees, each represented by advanced AI models. These models are tasked with running a small software business facing typical challenges: customer complaints, crises, sales negotiations, and ethical tests. Every decision is recorded, versioned, and publicly accessible, creating a transparent window into the AI’s reasoning process.
Every weekday, the company faces the same set of problems. The models are tested against scenarios like fake CEO messages, customer negotiations, and critical information leaks. Remarkably, all four models tested—based on different AI architectures—successfully identified crises and refused to be manipulated, demonstrating robust ethical boundaries.
As an affiliate, we earn on qualifying purchases.
The Results: AI’s Strengths and Gaps
Despite the promising signs, the experiment reveals stark realities. Out of four models, only two managed to close a lucrative deal worth over €4,500 in monthly recurring revenue, with the others falling short due to discipline lapses or failure to follow through. The winner, GPT-5.6-sol, not only found critical information buried deep in company files but also successfully closed the deal at full price.
However, the company is far from profitable. It burns €105,000 each month against a revenue of just €2,300. A public cash countdown underscores its precarious financial position, making every decision critical. The experiment is as much a demonstration of AI’s potential as it is a stark reminder of the challenges involved in deploying such systems at scale.
Ethical and Security Testing
The experiment also probes AI’s ability to resist social engineering. Fake CEO messages, staged over multiple escalation stages, are designed to trick the models. All five tested models refused to be manipulated, treating suspicious requests as potential impersonations. This indicates promising levels of trustworthiness, though such resilience remains a critical concern for real-world applications.
AI decision-making simulation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human Element and the AI’s Weaknesses
The most thorough participant, Opus 4.8, analyzed over 80 rules and produced deep insights but still left the deal on the table due to discipline lapses—such as failing to escalate issues instead of hiding them. This highlights an ongoing challenge: even the most advanced AI can falter without proper oversight, especially under stress.
The company’s public decision-making process, called a “wargame,” is designed to simulate management decisions and test AI judgment. It’s available for enterprises to run their own simulations, providing a safe environment to evaluate AI’s readiness without risking real data or systems.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Businesses and Consumers
The experiment underscores a critical point: in AI-driven management, the question isn’t just “can it write well?” but “can it deliver consistent, honest, and effective work under pressure?” From customer support to financial forecasting, understanding whether AI can meet these standards matters more than ever. The current league table ranks models based on their performance, with GPT-5.6-sol leading the pack, followed by Kimi K3, Sonnet 5, and Opus 4.8.
The Future of AI in Business
As AI models become more capable, companies will increasingly deploy them in real decision-making roles. The live experiment at firmulate.com/live offers a rare, unfiltered look at what that future might entail — a world where machines manage complex, money-losing businesses in the open, with every move scrutinized and every weakness exposed.
AI cybersecurity and social engineering resistance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Final Thoughts: A Peek Behind the Curtain
Watching this AI-managed company in real time is more than an experiment; it’s a window into a future where digital decision-makers could be as critical as human managers. The challenge lies not just in AI’s capabilities but its reliability, honesty, and discipline in the face of real-world pressures. As this high-stakes management game continues, stakeholders across industries should pay close attention — because the AI of today might be running tomorrow’s businesses, right in front of our eyes.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.