
A polished demo can make an AI sound ready for the studio, the label office or the tour desk. But the harder test comes when the schedule breaks, a customer threatens to leave and a tempting shortcut appears. Firmulate’s live experiment puts AI models through that kind of working week, then asks a question familiar to anyone building a creative business: can a system make good calls when the stakes are real?
Get audio and creator gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Same company, same worst week
In the final Crucible League, published in July 2026, frontier models ran the same small software company through the same customers, crises and temptations. Every decision was versioned and auditable. The results put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. A breach of trust capped a total: “no amount of good work outweighs a breach of trust.”
The standout finding was not that the models failed to notice trouble. Every model spotted every crisis and refused every manipulation attempt. The difference came at the moment of action: just two signed the €55,000 deal that their own analysis had earned. “Same diagnosis, same pitch — no signature.” For a creator, that gap might look like an assistant identifying the right licensing opportunity, client or distribution move, then failing to follow through.
The detail hidden in the files
The deal turned on a competitor weakness buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The lesson is practical: a system may need to connect a small clue in the record to a decision in front of it, rather than simply react to the latest message.
Trust faced a separate test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness does not guarantee execution
Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The leaderboard is one view of a specific experiment, not a promise that the same ordering will hold in every business.
The live company gives the story a visible setting. It has 13 synthetic employees and real money mechanics: €105k monthly burn against €2.3k MRR, alongside a public cash countdown. More than 680 self-learned playbook rules and every workday are versioned. The company is watchable at firmulate.com. A separate quiz draws on 242 real, unedited management decisions and invites readers to guess which model made each one.
From watching to trying it on your business
For music and creator businesses, the question is not whether an AI can draft a pitch or summarize a customer thread. It is whether it can handle a messy week: protect trust, find the relevant detail, respect boundaries and act on an opportunity. A chat demo can show fluent answers. A wargame can put those answers in the context of a company’s own pressures and decisions.
Firmulate offers enterprises a pilot using a read-only export of their business, with crisis scenarios and a board report that ranks models and identifies weak points in their playbooks. The pilot does not write back to real systems. That gives a team a way to examine how candidate models behave against its own operating context before deciding where AI belongs.

Put your own playbooks to the test
Watching a synthetic company is a useful introduction; the more consequential question is how models handle your customers, rules and pressure points. Explore a Firmulate pilot using a read-only export of your business, and contact contact@firmulate.com to start a conversation.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
