---
title: AI Decision Case Studies, Scored Against Reality | 3Dogs Nexus
description: Fourteen published case studies: a 10,000-page M&A data room, 45,320 Enron emails read blind, calibrated wildfire forecasting, and two where rival AIs reviewed our output verbatim.
url: https://3dogs.ai/case-studies/
---

- 
AI Decision Case Studies, Scored Against Reality | 3Dogs Nexus

Start a Decision Case →

 🐾3Dogs NexusStructured Decision Intelligence

 Home
 YouTube
 Start a Decision Case

 Case studies

# Real decisions. Shown work. Scored against reality.

Where can I see real examples of multi-model AI decision analysis?

Here. Every case below was run end-to-end and published in full, including the run statistics, the panel vote, and the dissenting views. Two of them were driven and reviewed by rival AI models — Google’s Gemini and OpenAI’s ChatGPT — with their criticism printed verbatim. These are internal and adversarial stress tests, not paid client engagements, and we say so on every page.

 We run actual questions through 3Dogs Nexus, put our reasoning on the table, and — where we can — get graded against the outcome. Here’s where that lives.

 Systemic risk · authored by a rival AI

### The AI economy's biggest bet, examined by 32 AIs at once

 Google's Gemini wrote the hardest question in tech finance — Oracle's BBB- rating, ~50% OpenAI concentration in a $638B backlog, and a "the government would never let it fail" theory — and was told to keep it politically neutral. Nine adversarial panels later, the answer came back as odds, not adjectives: ~15% chance of a formal bailout even if the deal fails, ~55% that help arrives quietly through procurement and national-security channels, ~60% that the media empire is already the hedge, ~8% bankruptcy. Both camps in the public debate are probably wrong. Not investment advice — a falsifiable forecast, dissent printed.

 Panel16 of 16, dissent kept

 3Dogs2,036 calls · 57m 00s

 32 models · 3 clouds9 debate panelsEvidence-classified

 Read the case study →

 Language stress test · production case

### Greek and Mandarin to open. Pitjantjatjara to clarify. French to finish.

 A real €12B case arrived in four languages across three messages, with zero coordination and no translation step. 3Dogs Nexus extracted a clean, internally consistent English brief — provenance-tracked back to the exact original-language sentence — then ran a 13-model debate to one decisive, unanimous call.

 Panel12 of 12, unanimous

 3Dogs201 calls · 5m 58s

 4 languages · 1 case13 modelsProvenance-verified

 Read the case study →

 Research integrity · reviewed by rivals

### The Devil's Advocate killed its own mission

 MIT disavowed a viral AI research paper that had already reached the European Central Bank and Congress. Google's Gemini turned the wreckage into a mission for 3Dogs Nexus — and one of our own 18 analysts flipped mid-debate to REJECT, arguing a volunteer AI panel had no business trying to audit an institution. We asked our Executive Committee if we should write the paper it should have been. It said no, and told us why.

 Panel1-15-1-1, dissent kept

 3Dogs1,565 calls · 30m 49s

 Research integrity29 models · 3 cloudsA REJECT vote, preserved

 Read the case study →

 Government AI assurance · reviewed by rivals

### The $97,000 lesson: we ran Deloitte's own AI failure through 3Dogs Nexus

 Deloitte refunded the Australian government after a single-model AI report fabricated a Federal Court quote. Google's Gemini turned that public failure into a benchmark — 18 adversarial analysts, 22 AI models, and a panel that names its own worst-case risk (“consensual hallucination”) before anyone else can.

 Panel18 of 18, with conditions

 Deloitte's refundAU$97,000

 Government & compliance22 models · 745 callsDissent named, not buried

 Read the case study →

 2008 crisis · head-to-head vs. Warren Buffett

### What if Lehman Brothers had "sold it all"? The Margin Call decision, run through 50+ AIs

 Two 2008 Lehman decisions, handed to 3Dogs. Should an investor have saved it? Buffett passed — and 3Dogs matched him with a 10-to-1 "walk away" that named Repo 105 and the 30-to-1 leverage. Could a Margin Call fire-sale have worked? On that unknowable counterfactual it stayed honestly uncertain. Calibration, not hindsight.

 Buffett3Dogs matched: REJECT

 CounterfactualLOW confidence · 5–3 split

 Warren Buffett head-to-headMargin Call counterfactualRepo 105 · 30:1 leverage

 Read the case study →

 Critical infrastructure · ransomware

### Pay the $18M ransom, or rebuild? 23 AIs split 10-to-9

 A water utility serving 420,000 people is locked out of its control systems by ransomware, with four days to pay. We put it to 23 AI models across AWS, Azure and Google — they split almost evenly, then landed on refuse and rebuild, with the strong pay-minority named, not buried. The panel’s sharpest insight: manual operations are both the fix and the real risk.

 VerdictREFUSE · rebuild

 Panel10–9 · 1,036 calls

 Public safety23 models · 3 clouds18 minutes

 Read the case study →

 Investigation · blind test on a real fraud

### The half-million-email test: can AI find the fraud buried in 517,401 emails?

 We handed 3Dogs the real Enron executive email archive with no hint of what to look for. It read 45,320 emails, surfaced the LJM, Raptor and Chewco off-books schemes and the overridden warnings from Kaminski and Watkins, and called for a formal investigation — in 2h 28m, versus roughly 4.5 years for the real investigation. The report priced the human review of this volume at $2–5 million.

 Read in full45,320 emails

 3Dogs2h 28m · read in full

 Early case assessment517,401-email corpusvs. $2–5M forensic review

 Read the case study →

 Document-scale · 8 / 8 recall

### The 10,000-page test: can any AI actually do due diligence?

 We built a 100-document acquisition data room, hid eight deal-killers inside, and put it to every option a buyer really has. Consumer AIs can’t even ingest it. A firm bills six weeks and five figures. The new “AI board” tools aren’t built for it. 3Dogs read all 10,000 pages, found all 8 buried risks, and returned a decisive RENEGOTIATE — in 28 minutes.

 Buried risks8 / 8 found

 Runtime28m 04s

 M&A due diligence100 docs · ~10,000 pagesvs. AI, firms & rivals

 Read the case study →

 Adversarial test · reviewed by a rival AI

### Google’s Gemini tried to break the system — then wrote the review

 Gemini played a hostile client: thin data, a napkin $3.5M estimate, critical context withheld on purpose. 3Dogs refused to analyze until three rounds of questioning — then Gemini published its verdict, word for word: “a hostile board member” producing “more rigor in 35 minutes than most human teams could produce in a week of meetings.”

 Geminihostile client

 3Dogs1,127 calls · 35m 40s

 SaaS rewrite · $4M runway17 modelsUnedited AI review

 Read the case study →

 Head-to-head · falsifiable

### Gemini vs. 3Dogs: the Lincoln County wildfire forecast

 Same question to both. Google Gemini answered in seconds — and did it well. 3Dogs answered in 4 minutes with a probability spread, a 9-analyst debate, graded evidence, and a date to be scored on. The right tool depends on who’s asking.

 Gemini~seconds

 3Dogs4m 20s

 PredictionPublic safety / planningResolves Oct 1, 2026

 Read the case study →

 Inside deep mode · full run log

### 18 minutes, two full ensembles, one decisive call

 We forced deep mode on a hospital capital-vs-contract decision and published everything: four adaptive research passes, a real clarification loop, six independently-composed debate panels, a live 429 failover, and a moment where the system caught a flaw in its own report.

 Ensembles6 panels

 Runtime18m 02s

 Capital investmentHealthcareDeep mode — forced

 Read the case study →

 Four decisions · zero hindsight

### 3Dogs vs. the war college vs. history

 We fed 3Dogs four of the most-taught command decisions in military history — Cuba 1962, Inchon 1950, Gettysburg 1863, Task Force Smith 1950 — using only what the commander knew at the time. Three times it matched the historical decision. Once it disagreed with the commander and matched the war college's own century-old critique instead.

 Matched3 of 4

 Diverged1 of 4

 Military strategyHistorical calibration test4 full runs

 Read the case study →

 Public-sector decision support

### Ranking a state DOT's road-safety projects — with the audit trail

 A 12-model panel prioritized safety investments by injury reduction, cost, readiness, equity, and community support — and named the exact conditions (independent score audits, Bayesian crash-factor adjustment) that make the ranking defensible to a legislature.

 Transportation11–0 proceed-with-conditions161 calls · 5m 46s

 Read the case study →

 Economic development

### Should a city fund a grocery feasibility study? The panel said: not first.

 Instead of rubber-stamping a $50K–$75K study, the 12-model panel reordered the sequence — cheap baseline first, then a study co-invested with grocery operators who commit in writing. The dominant risk it named: operator appetite, not demand.

 Market feasibility11–0 proceed-with-conditions156 calls · 2m 47s

 Read the case study →

 More on the way.Hiring calls, market entries, build-vs-buy, capital projects — real decisions, same treatment. Growing from here.

## Why we publish these

 Anyone can sound confident. Calibration is the only edge that compounds — so we show the reasoning, surface the disagreement, and put an expiration date on our forecasts. If we’re going to claim a better process, the honest thing is to let you watch it work and check it against what actually happened.

 “We don’t make your decisions. We make them better.”

 3Dogs Nexus · Structured Decision Intelligence · 3dogs.ai

 Terms · Privacy · Security

## Questions this case answers

### Are these real client engagements?

No, and we disclose that everywhere. They are internal and adversarial stress tests run on real public data — the actual Enron corpus, real historical decisions, real published financials. Using public datasets means anyone can check our work, which is the point.

### Have any been reviewed independently?

Yes. Google’s Gemini and OpenAI’s ChatGPT each ran a full case and wrote a review. Both are published unedited, criticism included.

### Try this on your own question.

 Free, no card. Bring a real decision — ideally one where you already know the answer —
 and see what the panel does with it.

 Start a decision case
