
Imagine managing a small greenhouse or outdoor living store where every decision counts, but instead of human staff, AI models are making the calls—facing crises, temptations, and the relentless pressure of real money. Now, picture watching that business struggle every single day, live online, as it tries to stay afloat. This is not science fiction. It’s the live experiment conducted by Firmulate, a company that runs AI-powered businesses in real time, revealing what it really takes for AI to act like a trustworthy manager in high-stakes situations.
The Experiment: AI as a Business Manager in the Wild
At the heart of this groundbreaking project lies a small, real software company operated entirely by AI models. Every weekday, the company faces typical crises—customer complaints, internal miscommunications, and financial pressures—just as a small outdoor gear or garden supply store might. But here’s the twist: instead of employees, 13 synthetic ‘employees’ make decisions based on a self-learned playbook of over 680 rules, updated daily and fully auditable.
This setup isn’t just a simulation; it’s a transparent, competitive experiment. Four frontier AI models—each with different strengths—are tested against the same week of chaos. The goal? To see which model can best manage the business, solve crises, and, crucially, close sales that sustain the company.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Models Achieved
All four models identified every crisis, from customer complaints to systemic issues, and refused manipulation attempts designed to trick them—like fake CEO messages or background approval requests. This shows they can recognize threats and maintain integrity under pressure, a critical competency for trustworthiness.
However, the differences lie in their ability to act decisively and follow through. Only two models managed to close the €55,000 deal their own analysis had earned, despite diagnosing every problem correctly and making the same pitches. The other two, including one with the most thorough analysis, left the deal on the table due to slipping discipline—failing to escalate issues properly or to act swiftly.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: What Lies in the Files
Interestingly, the decisive advantage came from reading deeper into the company’s own documents—specifically, a buried reference within their files that revealed an opportunity or weakness not immediately visible from customer interactions alone. Models that accessed these internal documents secured the full-paying deal, adding €4,583 in monthly recurring revenue.

The Operational Excellence Library; Mastering AI-Powered Chatbots in Customer Service
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Real Money, Real Stakes, Live Watching
The entire operation runs in real time at firmulate.com/live.html. It burns €105,000 every month against a modest €2,300 monthly recurring revenue, making it a microcosm of real business struggles. Every decision, every rule learned, and each crisis managed is visible to the public, offering unprecedented insights into how AI agents behave in complex, pressure-filled environments.

MASTERING CORPORATE FINANCE WITH CLAUDE AI: An Independent Guide to Financial Analysis, Forecasting, Automation, and Decision-Making
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Learning from the AI’s Performance
The experiment also included a detailed scoring system, where GPT-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, and Opus 4.8 scored 77. The top scorer, GPT-5.6-sol, was able to find the hidden opportunity that clinched the deal, showing how critical deep document analysis is. Meanwhile, Opus 4.8, despite having the most thorough rules, left the close on the table due to discipline lapses—a reminder that thoroughness alone isn’t enough without decisive action.
Trust and Integrity in AI Decision-Making
A noteworthy aspect of the experiment was the AI’s resistance to social engineering attempts. Fake CEO messages designed to bypass approval or impersonate leadership were refused across all models. Kimi K3 explicitly recognized these as suspicious, demonstrating an ability to assess threats to its trustworthiness.
Implications for Outdoor and Garden Businesses
While this experiment centers on a software company, the lessons resonate broadly—even for garden centers and outdoor living stores considering AI tools for customer service, inventory management, or sales. The key questions aren’t just about AI’s ability to generate convincing chat responses, but whether it can reliably follow through, read critical internal documents, and uphold trust under pressure. These qualities determine whether AI can become a dependable partner or just a flashy prototype.
The Build-in-Public Approach: Transparency as a Lens
What makes Firmulate’s live experiment especially compelling is its transparency. Every decision, every crisis, and every outcome is public and traceable. This build-in-public approach provides clear benchmarks and lessons—not only about AI’s current capabilities but also about what it takes to deploy trustworthy AI in real business scenarios.
Conclusion: Watching AI’s Business Journey Unfold
For anyone in the outdoor or garden space curious about AI’s future, this live window offers a rare glimpse into what’s possible—and what remains challenging. The experiment underscores that successful AI management involves more than clever chat. It demands discipline, deep understanding, and an unwavering commitment to honesty—traits that are still being tested in real-time, live environments. As these AI models continue to learn and improve, the question isn’t just whether they’ll write well, but if they can really run a business—trustworthy, resilient, and ready for the unexpected.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html