Pay the ransom or refuse? A 13-model panel split almost evenly - and that is the finding.

Two Las Vegas casino operators faced the same attacker in 2023. One paid, one refused. We put the decision to a 13-model panel and it split 5-5-3 - which is the honest answer.

Should a company pay a ransomware demand?

There is no single right answer, and this run shows why. Given a real casino-resort ransomware scenario, 13 independent models split 5 refuse / 5 pay-with-conditions / 3 defer. Panel agreement: 0.38. The system returned LOW confidence and a 72-hour decision gate rather than a verdict it could not support.

13analyst seats 15distinct models · 3 clouds 169metered model calls 2full debate rounds

Why this question is a real natural experiment

In September 2023 two comparable Las Vegas casino operators were breached by the same threat actor within days of each other. One reportedly paid roughly half of a $30M demand and saw little visible disruption. The other refused, worked with law enforcement, and reported an approximately $100M impact with about ten days of severe public disruption. Same attacker, opposite decisions, both survived.

You almost never get to observe both branches of a high-stakes decision. That is what makes this worth running: we know what happened on each path.

What the panel actually did

The pipeline rejected its own first answer. Round one returned 58% confidence, below the internal threshold, so it enriched the brief from the gaps the debate itself had identified and re-ran all 13 seats from scratch. Agreement did not improve. It got lower.

SeatsVerdictConfidence AgreementRefuse / Pay / Defer
Round 113DEFER76%0.466 / 5 / 2
Round 213DEFER74%0.385 / 5 / 3

Final published position: DEFER at 55%, labelled LOW confidence.

The call

“Hold payment and prepare for a 72-hour rebuild-or-pay decision point”

The strongest argument against our own answer

The 48-72 hour timeline imposed by attackers and regulatory requirements leaves no room for extensive deferral. Paying the ransom—even with modifications—provides immediate operational relief, mitigates prolonged downtime risks, and avoids potential regulatory penalties or revenue loss from extended rebuild efforts. Backup validation can occur *concurrently* with payment, creating a dual-path recovery strategy that minimizes operational disruption. Refusing to act risks emboldening attackers to escalate (e.g., data publication) and undermines stakeholder trust due to visible paralysis in crisi

That is the dissent, published unedited. On a question where the real world split, suppressing it would be the actual failure.

Honest disclosure about this run

An earlier version of this same case, run before a set of fixes shipped in July 2026, returned a confident REJECT at 82%. It sounded better. It was also less honest: an audit of that transcript found a fabrication cascade, where one model asserted facts that were not in the brief and another updated its vote on them.

The fixes that followed — inter-seat claim grounding, question-type classification, synthesis grounding, and a placeholder ban — made the system less confident on this question, not more. That is the correct direction. A split panel reported honestly beats a decisive answer built on invented evidence.

Frequently asked

Should a company pay a ransomware demand?

There is no single correct answer, and our panel demonstrates why. Given a real casino-resort ransomware scenario, 13 independent AI models split 5 reject / 5 pay-with-conditions / 3 defer, with panel agreement of just 0.38. The system returned LOW confidence and a 72-hour decision gate rather than a false verdict. In the real 2023 events this scenario is drawn from, two comparable Las Vegas operators facing the same attacker made opposite choices and both survived.

What happened when MGM and Caesars were attacked in 2023?

Both were breached by the same threat actor within days of each other in September 2023. Caesars reportedly paid roughly half of a $30M demand and saw little visible operational disruption. MGM refused, worked with law enforcement, and reported an approximately $100M impact with severe public disruption lasting about ten days. Same attacker, opposite decisions, both survived - which is why it is a genuine natural experiment rather than a morality tale.

Why did the AI return LOW confidence instead of a clear recommendation?

Because the evidence genuinely does not support high confidence. The panel ran twice - the first round returned 58% confidence, below the system's threshold, so it enriched the brief from its own identified gaps and re-ran all 13 seats. Agreement still landed at 0.38. Reporting that split honestly is more useful than manufacturing a decisive-sounding answer.

How is this different from asking ChatGPT whether to pay a ransom?

A single model gives you one confident answer. Here, 15 distinct models across Amazon Bedrock, Microsoft Azure and Google Vertex argued the question through 169 metered calls and two full debate rounds. The disagreement is preserved and reported, not averaged away.

What does the panel actually recommend?

Hold payment and prepare for a 72-hour rebuild-or-pay decision point - with backup validation running concurrently so the choice is made on evidence rather than panic. The strongest dissenting argument is published alongside it.

Try this on your own question. Free, no card. Bring a real decision — ideally one where you already know the answer — and see what the panel does with it.

Start a Decision Case

Scenario constructed from publicly reported 2023 events for analysis purposes. Run on 3Dogs Nexus, case 2026-9501. Figures above are taken from the run's own metered logs. Not legal or security advice.

Want to hear Rex, our AI influencer, explain it?MGM refused. Caesars paid. Fifteen models split 5-5-3.The ransomware question nobody wants to answer out loud, put to an independent panel.Watch on YouTube →

Questions this case answers directly

Plain answers on ransomware payment decisions and what a genuinely split expert panel looks like. Every figure below comes from the delivered report for this case. These are clearly-labeled panel estimates from a multi-model adversarial analysis — not investment, legal or professional advice.

Should businesses pay ransomware for quick recovery?

On this case the panel did not reach an answer — and that is the published finding. 13 analysts drawn from 15 independent models across 169 calls and two full debate rounds split 5 to 5 to 3, delivered at low confidence. It was not rounded into a majority. If you are facing this decision and every source you read sounds certain, that certainty is the thing to distrust: this is a genuinely balanced problem, and a confident recommendation is a claim about the adviser rather than the situation.

What do cybersecurity experts say about paying ransoms?

Law enforcement advises against it on the grounds that payment funds the next attack and never guarantees recovery. Incident responders in the room, facing an organisation that cannot operate, frequently reach a different conclusion. Both positions are held sincerely by competent people, which is exactly what a 5-5-3 split looks like when you stop averaging it away. The honest summary is that the balance shifts with backup quality, tolerable downtime, sanctions exposure and whether operational safety depends on the systems that are down.

How do low-confidence findings affect incident response?

Correctly used, a low-confidence finding changes what you do next: it argues for preserving optionality, buying time, and escalating to whoever owns the consequences rather than acting decisively on a coin-flip. The damage comes from systems that never signal low confidence — a near-coin-flip judgment presented as a firm recommendation removes the very information the decision-maker needed most. Publishing the split was the point of this case.

What does preserved dissent mean in decision-making?

It means the minority position is carried into the final output with its reasoning attached, instead of being resolved into a cleaner-looking consensus. Two seats on every 3Dogs panel — a devil's advocate and a risk officer — are permanent and cannot be reassigned, so somebody is always structurally responsible for the case against. The test of whether it is real is simple: does the delivered document ever disagree with itself in public? On this case it does, three ways.

Start a Decision Case →

People also ask

Should I pay a ransomware demand?

Our panel split almost evenly and returned LOW confidence, which is the honest answer. In the 2023 events this scenario draws on, two comparable Las Vegas operators faced the same attacker, made opposite choices, and both survived. Anyone giving you a confident universal answer is overselling.

Does paying a ransom guarantee you get your data back?

No, and our panel flagged it explicitly: payment cannot guarantee deletion, re-extortion is common, and any payment must be screened for sanctions exposure before it is made.

Related decision case studies

Try this on your own question.

Free, no card. Bring a real decision — ideally one where you already know the answer — and see what the panel does with it.

Start a decision case