The Management Test That Discloses AI’s Genuine Working Style
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The Management Test That Discloses AI’s Genuine Working Style on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get everyday essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A new management test, conducted by Firmulate, evaluates AI models on a simulated company’s worst week. The experiment reveals distinct differences in how models identify problems, act decisively, and maintain trust. This offers insights into AI’s practical management capabilities.

Firmulate has launched a live management test involving five AI models managing a simulated software company’s worst week, revealing clear differences in their ability to follow through on decisions, protect trust, and complete critical actions. This experiment provides a rare, observable window into how AI handles complex management tasks under pressure, making it highly relevant for enterprises considering AI management automation.

The experiment involves five AI managers operating a simulated company with 13 synthetic employees, a monthly burn rate of €105,000, and €2,300 in recurring revenue. For more on evaluating AI models, see the original analysis. Each AI faces the same crises, customer scenarios, and decision points, with their actions recorded and auditable. The models are scored based on their ability to diagnose problems, act decisively, and avoid breaches of trust. The top performer, gpt-5.6-sol, scored 95 points out of 100, while others like Kimi K3 and Sonnet 5 followed closely behind. A baseline model scored 26, highlighting the gap between minimal effort and effective management.

One key finding is that all models recognized crises and refused manipulation attempts, such as fake CEO messages. This aligns with the insights from the detailed management test analysis. However, only two models successfully closed a crucial €55,000 deal, demonstrating that identifying issues is not enough—executing the right actions is essential. The experiment also revealed that thorough analysis alone does not guarantee success; models like Opus 4.8, despite extensive reasoning, failed to complete operational tasks, such as escalating issues or closing deals.

At a glance
reportWhen: ongoing; results published July 2026
The developmentFirmulate’s live experiment tests five AI management models on a simulated company’s crisis week, revealing their strengths and weaknesses in decision-making and trust management.

Implications for AI Management Effectiveness

This experiment underscores that effective AI management requires more than analysis and recognition of problems. Successful models must also execute decisive actions, maintain trust, and follow through on operational steps. For businesses, this highlights the importance of testing AI systems in realistic, pressure-filled scenarios before deploying them in real-world management roles. The results demonstrate that AI’s ability to act reliably and decisively is crucial for operational success and trustworthiness in automation.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Management Testing

Traditional AI demonstrations often focus on analysis and problem identification, but real-world management involves execution, trust, and follow-through. Firmulate’s live experiment builds on prior efforts to evaluate AI in operational settings by creating a simulated environment where models face realistic crises and decision-making pressures. The league results from July 2026 compare five frontier models, with the highest scoring at 95 points, illustrating the current state of AI management capabilities and limitations.

“The experiment reveals that recognizing crises is not enough; effective management requires decisive action and trust preservation.”

— source from Firmulate

Amazon

business AI simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Aspects of AI Performance Are Still Unclear

It is not yet clear how these results will translate to real-world management settings, where variables and stakes may differ. The experiment’s simulated environment, while realistic, cannot fully replicate the complexities of live business operations. Additionally, the long-term reliability of these models in ongoing management tasks remains to be evaluated as AI systems evolve and are further tested in diverse scenarios.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Steps for AI Management Evaluation

Further testing will focus on deploying these models in actual business environments, observing their ability to handle ongoing operations over extended periods. Firms may also develop customized assessments based on their specific workflows and risks. The results from these experiments are expected to inform best practices for AI management, including how to measure and improve operational follow-through and trustworthiness.

Amazon

AI trust and decision execution tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main purpose of Firmulate’s management test?

The test aims to evaluate how AI models handle realistic management scenarios, focusing on decision-making, execution, and trust preservation under pressure.

How are the AI models scored in the experiment?

Models are scored based on their ability to diagnose crises, act decisively, avoid breaches of trust, and complete critical operational tasks, with the highest scorer receiving 95 points out of 100.

What does the experiment reveal about AI’s management capabilities?

It shows that while AI can recognize problems and refuse manipulation, effective management also requires decisive action and follow-through, which not all models currently perform well.

Will these results be applicable to real-world businesses?

While the experiment provides valuable insights, real-world applicability depends on further testing in actual operational environments to verify AI’s reliability and effectiveness over time.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The deployment. How the AI labs verticallyintegrated into the serviceslayer — the Palantir modelat scale.

Major AI labs deploy forward-engineering models to embed AI into enterprise operations, reshaping the consulting industry and revenue streams.

The Future Of Corporate Wellness: Remote Work Strength Tests

New remote work strength assessments using phone cameras aim to detect early musculoskeletal issues, offering a preventive approach for employers.

A Skill Is a Folder, Not a Prompt: What Anthropic Learned Running Hundreds of Them

Anthropic reveals that Skills are folders containing instructions, scripts, and assets, not just prompts, enabling more durable AI capabilities.

The Role Of Virtual Reality In Modern Manufacturing Workforce Prep

Manufacturers are testing VR modules for faster, safer worker onboarding amid labor shortages. Pilot programs show promising results in safety and efficiency.