
Imagine a travel company hit by a sudden wave of cancellations, a competitor undercutting its prices and a fake message from the CEO asking for a quick exception. Would an AI assistant spot the trouble? And could it follow through on the right response without crossing a line?
Get travel and outdoor gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Firmulate put AI models through that kind of pressure in a live experiment. The company is synthetic, but the business stakes are designed to feel familiar to anyone who runs bookings, customer support or an outdoor travel operation: protect trust, respond to a crisis and make the sale when the evidence supports it.
Same company, same rough week
In the final Crucible League, published in July 2026, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were held constant, and every decision was versioned and auditable. The point was to examine management quality under pressure, not just how polished a model sounds in a chat window.
The models all recognized every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.” A system can understand what should happen and still fail to complete the job.
The final ranking placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The detail hidden in the company’s files
The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a practical reminder for travel businesses: useful context may be sitting in existing notes and records, well away from the latest booking or support request.
The pressure tests also included fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” approach. All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of caution matters in any company handling customer information, urgent exceptions or confidential plans.
Thoroughness is not the same as follow-through
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and discipline slipped: it made write attempts into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
There is a fairness detail to keep in view: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context belongs alongside the ranking when readers consider what the comparison can show.
Firmulate’s live company gives the experiment an ongoing, watchable dimension. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The site also offers a “guess the model” quiz built from 242 real, unedited management decisions.
Readers can follow the live experiment at firmulate.com. The experiment is a public demonstration of models managing a company; its lesson for operators is that recognizing a problem, protecting trust and carrying a sound decision through are separate challenges.
From watching to trying it on your own business
For a travel or outdoor company, the useful question is not simply whether an AI can answer a customer. It is how the system behaves when cancellations spike, a competitor makes a move, a sensitive request arrives or an opportunity depends on context buried in company records.
Enterprises can run the same kind of wargame against a read-only export of their own business. The pilot tests crisis scenarios against that company context and produces a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. Explore the Firmulate pilot or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
