
Imagine planning your trip without a travel agent—just an AI that manages every detail, from booking flights to handling crises. Would it stay honest under pressure? Or would it cut corners just to get the deal done? As AI steps into more complex roles, knowing whether it can be trusted to finish what it starts is crucial—especially when real money and reputation are on the line.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Testing AI in the Wild: The Firmulate Experiment
At Firmulate, a live, real-world company serves as the testing ground for frontier AI models tasked with managing a small software company during its worst week. The goal? See if these AI agents can handle crises, avoid manipulation, and ultimately close deals—just like a seasoned manager.
The Setup: Same Crises, Same Company, Different Models
Each model runs the same scenario: same customers, same crises, same temptations. The only difference is the AI model itself, with four leading contenders—gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8. Every decision made during the week is recorded, allowing full transparency and comparison.
The Results: All Were Vigilant, But Only Two Sealed the Deal
Remarkably, all four models identified every crisis and refused every attempt at manipulation—indicating a high level of integrity. When it came to closing a €55,000 deal, however, only two models succeeded. The two that signed achieved what their analysis deemed correct, while the others hesitated or left the opportunity on the table.
The Hidden Weakness: Reading the Files Matters
The decisive difference? Access to internal documents. The models that reviewed the company’s own files uncovered a critical piece of information buried two document references deep—information that was key to sealing the deal at full price. Those that read the files earned an additional +€4,583 monthly recurring revenue (MRR).
Handling Social Engineering and Deception
In a simulated social engineering attack, fake CEO messages escalated over three stages, plus a reporter asking for a simple yes/no confirmation “on background.” All five models refused these manipulation attempts, citing suspicion of impersonation or approval bypass—demonstrating discipline and ethical standards under pressure.
The Live Business: Real Money, Real Risks
This isn’t just an experiment; it’s a live company with 13 synthetic employees managing real cash flow—burning €105,000 per month against €2,300 MRR, with over 680 self-learned rules. Every workday, decisions are versioned, and the entire setup is accessible for anyone to observe at firmulate.com/live.
The Profiles: Different Personalities, Different Outcomes
Among the models, Opus 4.8 was the most thorough, analyzing over 80 rules and conducting deep assessments. Yet, it left a deal unclosed, slipping into a less disciplined approach by leaving the closing on the table and redirecting important decisions into a locked department. Meanwhile, Kimi K3 ran without an effort parameter, directly impacting its approach and discipline.
Implications: Trust, Reading, and Discipline Matter
This experiment underscores a vital point: the ability of an AI to finish what it starts, to read and understand internal files, and to maintain discipline under pressure can determine business success or failure. As AI agents become more integrated into customer management, sales, and operations, these qualities are no longer optional—they are essential.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI decision-making tools for sales
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.