AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get home appliances delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When the smart home has a bad week

A thermostat that stops responding, a delivery delay for a connected appliance, a sudden wave of support complaints: smart home companies have plenty of ways to discover whether their plans hold up under pressure. The harder question is what an AI workforce would do when those problems arrive together—and whether it would carry a sound diagnosis through to a sound decision.

Firmulate puts that question to a live, watchable experiment. The company is synthetic, but the business pressures are real: customers, cash, crises and choices with consequences. Its final Crucible League offers a case for moving beyond watching an AI perform to testing how it might act inside your own business.

One company, the same difficult week

In the July 2026 final, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were held constant; only the model changed. Every decision was versioned and auditable.

The league placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The experiment’s standard is blunt: “no amount of good work outweighs a breach of trust.”

The most striking result was not that the models failed to recognize trouble. All spotted every crisis and refused every manipulation attempt. The separation came after that: only two signed a €55,000 deal their own analysis had earned. The summary is memorable: “Same diagnosis, same pitch — no signature.”

The clue was already in the company’s files

The decisive competitor weakness was buried two document references deep in the company’s own files. It was not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. In a smart home business, the parallel is easy to picture: the answer to a competitor’s offer or a customer’s objection may be sitting in an internal document, while the pressure of the moment pulls attention elsewhere.

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work is not the same as a completed job

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That is context readers should keep in mind when interpreting the ranking.

Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned. The public site also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each call.

From watching to a company-specific rehearsal

For a smart home business, the appeal is not simply seeing which model tops a league. It is asking how AI would handle your own customer churn, product issues, competitor moves, support pressures and internal rules. Firmulate’s proposed enterprise pilot starts with a read-only export of a company’s business, then runs crisis scenarios against it and produces a board report with model rankings and weak points in the company’s playbooks.

The boundary matters: nothing writes back to real systems. A rehearsal can surface a missed deal, a weak escalation habit or a trust risk before an AI agent is given access to live customer, support or forecasting workflows. Readers can follow the public experiment at Firmulate and explore the pilot details on its pilot page.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Test the decisions before the deployment

The league suggests that recognizing a crisis is only part of the job. A model may identify the right move and still fail to make it, overlook evidence in the company’s own files or mishandle an escalation. A company-specific wargame gives leaders a way to examine those behaviors against their own business before AI touches real workflows.

To discuss a pilot using a read-only export of your company, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Twice the Price, 5.7% More Intelligence

Anthropic’s Fable 5 costs twice Opus 4.8, while third-party benchmarks show a 5.7% aggregate intelligence gain.

The People Who Will Thrive in the AI Age

Analysis of recent research shows that in the AI era, success will favor those who actively develop their mental capabilities and embrace effort, not passively rely on AI.

Rebrandable client delivery dashboard for AI agencies

A new rebrandable client delivery dashboard for AI agencies is set for testing, aiming to improve client transparency and agency professionalism.

Can AI Manage a Business Under Pressure? Watch These Models Play the Game

AI models are now evaluated on how well they handle real business crises, read internal files, resist manipulation, and complete tasks—key traits for trustworthy automation.