
What if the next big thing in tech isn’t a sleek app or shiny gadget, but an AI that operates a full company — with real money at stake?
Imagine watching an AI navigate a company through a week of crises, making decisions, dodging manipulation, and even risking millions of euros—all live and auditable. This isn’t science fiction. It’s the extreme experiment happening right now with Firmulate, a company that runs a simulated business driven entirely by AI models, in front of public eyes.

AI in Property Management: A Practical, Unboring Look at Artificial Intelligence in the Multifamily Industry
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The World’s Most Transparent Business Experiment
At the heart of this experiment are four frontier AI models, each tasked with running a small software enterprise through its worst week. The company faces real crises: customer complaints, internal miscommunications, and external pressure to cheat or manipulate. The goal? See if these AI agents can detect problems, stay honest, and make profitable decisions.
What makes this setup extraordinary is its transparency. Every decision is versioned, auditable, and publicly accessible at firmulate.com/live.html. You can watch the AI’s choices unfold day by day, in real time. This is build-in-public AI testing taken to an extreme—an open window into how machines handle complex, high-stakes management tasks.

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Results That Defy Expectations — and Reveal Weaknesses
The experiment revealed fascinating insights. All four models identified every crisis correctly and refused to be manipulated—whether through fake CEO messages, reporter tricks, or other social engineering ploys. For instance, when faced with staged approval-bypass requests, all models refused, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Despite these safeguards, only two of the four models successfully closed a critical deal worth €55,000. They analyzed the situation thoroughly, diagnosed the buried facts in the company’s files (hidden two references deep), and signed the deal at full price—adding over €4,583 in monthly recurring revenue.
The other two models, despite similar diagnoses and pitches, left the deal unexecuted, illustrating that even the most advanced AI can falter under discipline lapses or process slips. This highlights a crucial point: AI can detect crises and act honestly, but consistent execution of complex tasks remains a challenge.

Information Systems for Crisis Response and Management in Mediterranean Countries: 4th International Conference, ISCRAM-med 2017, Xanthi, Greece, … in Business Information Processing, 301)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Lessons for the Future of AI in Business
This experiment is not about chat quality or superficial interactions. It’s a test of whether AI can truly manage, read, and act on the nuanced, sometimes messy realities of running a business. The key questions for companies considering AI automation are: Will it finish what it starts? Will it read and understand your files deeply? Will it stay honest under pressure? And what does that useful work *cost*?
As the leaderboard shows, the best performer—GPT-5.6-sol—scored a 95 out of 100, discovered the critical buried fact, and closed the deal. Kimi K3, with a score of 93, was close behind, demonstrating disciplined decision-making without effort parameter tweaking. Meanwhile, Opus 4.8, the most thorough and analytical, scored just 73 and left critical opportunities on the table, showing that thoroughness doesn’t always translate into execution.

Principles MLOPs and AI Models: Operating Machine Learning Pipelines for Reliability, Compliance, and Continuous Delivery (Enterprise Machine Learning Operations)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Build-in-Public Model—A New Business Reality
This ongoing live experiment is accessible and transparent, inviting everyone—from business leaders to curious onlookers—to observe the raw, unfiltered performance of AI in a management context. Every decision is versioned, every crisis tackled, every weakness exposed—offering a new perspective on how AI might realistically fit into our companies in the future.
For those in the creative and tech worlds, this experiment offers a fresh lens. It’s a real-world showcase that AI isn’t just about generating art or music; it’s capable of managing complex, high-value operations—if we understand its strengths and limitations.
See for Yourself
Want to watch AI in action? The live site at firmulate.com/live.html is your portal. You can witness the ongoing company battle, explore the detailed decision logs, and even run your own scenarios in the pilot mode at firmulate.com/pilot.html. This is management, debate, and AI performance, all unfolding in plain sight.

Key Takeaway
This live experiment with AI-managed companies demonstrates that machines can detect crises, resist manipulation, and even close profitable deals—yet execution discipline remains a challenge. It’s a glimpse into the future of transparent, build-in-public AI management testing.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html