firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a music producer trying to perfect a track under pressure, knowing that every decision could make or break the hit. Now replace that producer with an AI managing a real company, facing the same crises—decisions, temptations, and pressure—and what you get is an eye-opening experiment in AI management. Just as a great producer stays disciplined amid chaos, some AI models are proving they can run a business with surprising integrity. The latest results from the live Firmulate experiment reveal that even newcomers can outperform established giants, reshaping how we think about deploying AI in real-world scenarios.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get audio and creator gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Real-World Test of AI Business Management

In July 2026, the Firmulate platform conducted a groundbreaking test—an experiment that put four leading AI models to the ultimate test: managing a small software company through its worst week. This wasn’t a simulation or a chat demo; it was a real, live, business-critical environment where every decision mattered, and every crisis had to be addressed with integrity.

The models faced the same customers, same crises, and the same temptations to cheat. Each decision was versioned and auditable, ensuring transparency in their choices. The goal? To see which AI could run the company most successfully, not just in terms of chat quality but in completing actual management tasks—reading files, making decisions, resisting manipulation, and ultimately closing deals.

Amazon

AI management decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Results That Defied Expectations

The standings in the final league table were telling:

  • gpt-5.6-sol scored the highest at 95, discovering critical buried information in the company files and closing a €55,000 deal, earning an additional €4,583 in monthly recurring revenue (MRR).
  • Kimi K3, the newcomer from Moonshot, scored 93, just behind, and demonstrated the cleanest discipline among all models, successfully resisting manipulation attempts and closing the same deal.
  • Sonnet 5 scored 88, also closing the deal but with more process slips, while Fable 5 and Opus 4.8 lagged behind at 77 and 73 respectively, with Opus leaving the close on the table and slipping in discipline.

Notably, all models spotted every crisis and refused manipulative tactics like fake CEO messages and reporter tricks. Only two models, gpt-5.6-sol and K3, signed the deal based on their own analysis. The others hesitated or failed to act at full potential.

Amazon

business AI simulation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Made the Difference?

A key insight was that the decisive weakness for some models wasn’t in their decision-making per se but in their ability to read and interpret critical internal documents—hidden references that contained vital facts. The models that successfully examined these files won the deal at full price, demonstrating that reading comprehension and thoroughness matter just as much as the final decision.

Another noteworthy aspect was the models’ resistance to social engineering. When fake messages from a CEO escalated through multiple stages, all five models refused to act on them without verification, citing suspicion of impersonation or approval bypass.

Amazon

AI decision-making tools for companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Company: A Real Testbed

The experiment was conducted on a real software company with 13 synthetic employees, operating under real cash mechanics—burning €105,000 monthly against a revenue of €2,300. Every workday, the company’s rules and decisions are logged and versioned, offering a transparent window into AI management decisions. You can watch these live operations at firmulate.com/live and see how different models perform in real-time.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Role of Discipline and Fairness

While most models demonstrated strong crisis detection and refusal to manipulate, the most thorough participant—Opus 4.8—showed discipline lapses and left opportunities unseized. Its weakness was linked to the tendency to escalate issues into locked departments rather than resolving them proactively. Interestingly, Kimi K3 ran at the default effort level (without an effort parameter), providing a fair comparison point against others that ran at the xhigh setting, highlighting how discipline influences outcomes.

The Broader Implication: Trust and Performance

This live experiment underscores a vital point for anyone deploying AI in business: it’s not just about how well the AI writes or chats, but whether it can complete work honestly and effectively under pressure. AI’s capability to read internal data, resist manipulation, and finish what it starts is now measurable—and critical.

As the league stands, the field is open. The newcomer, K3, beat three established models, proving that fresh entrants can outperform traditional players when rigorously tested in real-world scenarios. This shifts the landscape for enterprises considering AI solutions—picking a model without your own test is now a gamble.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The recent live experiment demonstrates that new AI models like Kimi K3 can outperform established leaders in managing real business crises, showing discipline, honesty, and thoroughness. Decision quality, not just chat prowess, now determines success—and enterprises should test their AI before deploying it at scale.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Odyssey Rotten Tomatoes

The latest Odyssey film has received a controversial Rotten Tomatoes score, prompting discussions among viewers and critics about its quality and reception.

Worst Sitcom You’ve Watched

A trending discussion on Bluesky has named the worst sitcom watched by users, sparking debate about comedy quality and viewer preferences.

Mickey Mouse Club Surges In Global Coverage

The Mickey Mouse Club is experiencing a surge in international media mentions, with 23 reports in a recent window, indicating heightened global interest.

Love Island Reunion

The Love Island reunion for 2026 has been officially announced, bringing together past contestants for a special event. Details are still emerging.