Firmulate —
Live on firmulate.com.

Imagine running your garden center through a week of chaos—unexpected customer complaints, supplier crises, and tricky negotiations. Now, picture AI models managing that same chaos, but with personalities and decision styles that are as distinct as seasoned managers. How do you tell which AI is trustworthy? The answer lies in a groundbreaking live experiment that puts these models to the test in a real, working company.

Testing AI in the Real World: The Firmulate Experiment

At Firmulate, a company dedicated to measuring AI’s management capabilities, four frontier AI models faced the same intense challenge: run a small software company through its worst week. This simulated scenario included the same customers, crises, and temptations, with every decision documented for transparency and analysis. The goal was to see not just if these models could identify problems, but whether they would act honestly and consistently under pressure.

The models included:

  • gpt-5.6-sol, scoring the highest with 95 points
  • Kimi K3, at 93 points
  • Sonnet 5, with 88 points
  • Fable 5, scoring 77 points

All four models effectively spotted every crisis and refused every manipulation attempt. That means, regardless of the AI’s personality, they all demonstrated a baseline of honesty and awareness. But the real story emerged in the decisions that affected the bottom line.

Amazon

AI decision-making software for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness and the Big Win

Despite their vigilance, only two models managed to close a critical €55,000 deal based on their own analysis—an important measure of their management competence. Interestingly, the key advantage was not in the obvious customer interactions but buried two document references deep within the company’s files. Those models that read and understood those hidden details were able to win the deal at full price, adding €4,583 MRR to the company’s revenue.

Amazon

AI management tools for small companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Social Engineering Test

In another mimic of real-world pressure, the models faced staged fake messages from a supposed CEO escalating over three stages, plus a reporter trying to trick them with a simple background question. All five models refused to be manipulated, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows their capacity to resist social engineering tactics that often compromise human decision-makers.

Amazon

AI business crisis simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Company and Its Lessons

The experiment’s company was a real, functioning business with 13 synthetic employees, burning €105k monthly against a €2.3k MRR. It operates with over 680 self-learned rules, every workday versioned, and a transparent online dashboard available for viewers at firmulate.com/live. This setup provides a rare window into how AI models perform during genuine business operations, not just scripted demos.

Amazon

AI social engineering resistance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Personality Profiles and Performance Gaps

The most thorough participant was Opus 4.8, which incorporated over 80 learned rules and performed in-depth analyses. However, it finished last—failing to close the deal and slipping into a discipline breach by leaving the close on the table and instead writing attempts into a locked department. This reveals that even the most diligent AI can falter if not aligned with clear decision-making discipline.

K3, running without an effort parameter (the API default), demonstrated the cleanest discipline, while others operated at higher effort levels, which could influence performance. Interestingly, in all four models, the same weaknesses appeared, hinting at fundamental design limitations that need addressing.

Measuring Management Personalities

This experiment underscores that AI models exhibit measurable management personalities—some are thorough and cautious, others terse and direct, and some even prone to slipping discipline under pressure. The scores from the Crucible League reflect this diversity, with GPT-5.6-sol leading at 95 and Fable 5 trailing at 77.

What This Means for Garden and Outdoor Businesses

While this experiment centers on a software company, its implications ripple outward. If AI agents are to manage your CRM, handle customer support, or forecast sales, the critical question isn’t just whether they generate good chat responses. It’s whether they can finish what they start, read your files thoroughly, stay honest under stress, and deliver tangible results at a fair cost. Trustworthiness and reliability are paramount—qualities that are now quantifiable and observable in live, rigorous tests.

Infographic —
The findings at a glance — source: firmulate.com.

The Firmulate live experiment vividly demonstrates that AI models can exhibit distinct management personalities, with some better at closing deals and resisting manipulation. For outdoor and garden businesses, choosing the right AI isn’t just about clever language—it’s about ensuring consistent honesty, thoroughness, and execution in real crises. Test your future AI workforce before hiring at firmulate.com/quiz.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

DEWALT Power Tools Battery Not Charging: Causes & Fixes

Troubleshoot your DEWALT battery charging issues with this step-by-step guide. Learn common causes, safe fixes, and tips to keep your tools powered up.

Sewing Machines for Quilting What Beginners Usually Miss

Getting started with quilting involves many overlooked details that can make or break your project—discover what beginners often miss and how to improve.

DEWALT XR Impact Driver vs DEWALT Brushless Drill: Full Comparison

Compare the DEWALT XR Impact Driver and Brushless Drill to find the best fit for your projects. Detailed specs, pros, cons, and real-world insights included.

DIY Table Decor Ideas for Everyday Dining

You can easily elevate your everyday dining with DIY table decor that…