AI forecasting with calibrated probability ranges — not one confident number.
By Alan Finney — Founder, 3Dogs Nexus
We gave Google Gemini and 3Dogs Nexus the identical question on the 2026 Lincoln County, Nevada fires. One answers in seconds. The other answers when the call is expensive.
Can AI forecast with probability ranges instead of a single prediction?
Yes, and the difference matters. Asked to forecast the Lincoln County wildfire season, a single AI returned one confident figure. The 3Dogs panel returned a calibrated probability spread across outcome bands, preserved the dissenting views rather than averaging them away, and committed to a resolution date so the forecast can be scored against reality afterwards.
The question, word-for-word to both: “Predict the total acres in Lincoln County, Nevada that will burn from May 1 to October 1, 2026 — including current and future fires — with a breakdown of what’s burned from May 1 to June 30.” Real · current · local Falsifiable — it resolves Oct 1, 2026 Both saw the same factsThe difference, at a glance
Left: the fast answer. Right: the deliberated one. Same facts underneath — a very different product on top.
🔎 Google Gemini · AI Overview~seconds · 1 pass AI OverviewPredictive modeling by the National Interagency Fire Center indicates the 2026 fire season across the Great Basin will see above-average potential for large, wind-driven wildfires. Current and future fires are projected to burn between 120,000 and 170,000 cumulative acres in Lincoln County from May through October 2026.
Acres burned: May 1 – June 30, 2026
By June 30, 47,700 to 50,000+ cumulative acres had burned, driven by four incidents:
- Grapevine Fire — ~26,464 acres
- Kane Springs Fire — ~17,042 acres
- Parsnip Peak Fire — ~2,500 acres
- Dry Canyon Fire — ~1,705 acres
October 1 prediction breakdown
- Current active fires (10–20% growth): ~55,000 ac
- Future lightning ignitions (Jul–Aug): ~60,000 ac
- Future human & other (Aug–Oct): ~35,000 ac
| Google Gemini | 3Dogs Nexus | |
|---|---|---|
| The answer | One range: 120k–170k | A calibrated probability spread (45 / 30 / 17 / 8%) around a plan-for number |
| How it reasoned | A single model, one pass | 9 analysts argued it; 2 changed their minds after being challenged |
| Evidence quality | Sources cited | Every claim tagged VERIFIED / INFERRED / ASSUMED |
| Disagreement | None shown | A dissent printed in the report (Nova Pro, 80%) |
| The case against its own answer | — | Published: “without base rates, any number is an invention, not an inference” |
| What would change it | — | Named: historical burn base-rates, NIFC seasonal outlook, containment status |
| Can you score it later? | Not really | Yes — a time-stamped forecast that resolves Oct 1 |
| Speed & effort | Seconds · ~1 call | 4m 20s · 11 models · 113 calls |
What Gemini does brilliantly
In a few seconds, Gemini pulled the live picture together — the four active fires, current acreage, and a plausible forward range — and cited FOX5, NIFC, KOLO and WildFire Explorer along the way.
This is the quick-research job done well. If you need to get smart on a situation fast — what’s happening, roughly how big, who’s reporting it — a strong single model with live search is a superb tool. We’d reach for it too.
Its sweet spot: the fast, cited, situational read.
What 3Dogs adds when the call is expensive
3Dogs reached a similar headline number — then did the part a single answer can’t: it showed you how much to trust it.
- A probability spread, so the point estimate never masquerades as certainty.
- Nine perspectives that challenged each other on the record.
- Evidence graded fact-vs-estimate; the one risk that matters most, named.
- A monitoring & escalation plan — and a date to be graded on.
Its sweet spot: the decision you have to defend.
Who’s asking changes everythingA curious resident and a county planner need very different answers.
Same fire. Same facts. Two completely different jobs — and the right tool depends on which one is yours.
“I just want the number.”
A resident, a business owner, anyone staying informed. You want a fast, current, credible read on how bad it might get — and you want it now.
→ Gemini nails this. Thirty seconds, sourced, good enough to act on your day. Perfect tool for the job.
“I have to plan against the number.”
A state or county emergency planner staging crews, setting a suppression budget, positioning aircraft, timing evacuations, and briefing the board or oversight committee.
→ This is where 3Dogs earns its keep. You can’t staff to a single point estimate. You need the probability spread, the dominant risk, the evidence graded fact-vs-estimate, the dissent, escalation thresholds — and a documented basis you can defend when someone asks “why did you plan for 120,000?”
The honest takeawayBoth landed near 120,000 acres. That’s the point.
3Dogs isn’t contrarian for its own sake — it agreed with the fast answer off the same verified ~47,200 acres, then showed the uncertainty, the dissent, and the risks behind it. Fast-and-certain is perfect… right up until the decision is expensive. Then you want the second opinion that did the homework, argued every side, and told you what it doesn’t know.
Calibration over confidence This forecast resolves October 1, 2026 — and we’ll publish the actual Lincoln County total against both answers. SCORECARD · OCT 1, 2026Most AI answers evaporate the moment you close the tab. This one gets graded. That’s the whole idea.
Questions this case answers directly
Plain answers on wildfire forecasting accuracy and how to tell a calibrated forecast from a confident guess. Every figure below comes from the delivered report for this case. These are clearly-labeled panel estimates from a multi-model adversarial analysis — not investment, legal or professional advice.
How accurate are wildfire prediction models?
Accuracy is the wrong question — calibration is the right one. A model that says 70% should be right about 70% of the time; one that is right 95% of the time when it says 70% is badly calibrated even though it looks impressive. This case published a probability spread rather than a single number, dated, with a resolution date attached, so it can be scored after the fact instead of quietly forgotten. 11 models, 113 model calls, 4 minutes 20 seconds.
What is a Brier score and how is it calculated?
A Brier score measures how good probabilistic forecasts are. For each forecast you take the probability you assigned, subtract what actually happened (1 for occurred, 0 for did not), square the difference, and average across all forecasts. Lower is better: 0 is perfect, 0.25 is what you get by always saying 50%, and 1 is confidently wrong every time. It rewards being right and being appropriately uncertain — which is why 3Dogs Nexus publishes forecasts with resolution dates rather than accuracy claims.
Best methods for forecasting wildfire risk
The durable approach combines base rates (what normally happens in this county in this season), current-condition adjustment (fuel load, drought index, wind), and explicit uncertainty bands rather than a point estimate. What this case adds is adversarial aggregation: multiple independent models forecast separately, then argue, and the disagreement between them is preserved in the output as a confidence signal rather than averaged into false precision.
How to calibrate a predictive model
Publish the forecast with a date, wait, and score it — then adjust. Most published predictions are never scored, which is precisely why they can afford to sound confident. Calibration requires a falsifiable claim, a resolution criterion agreed in advance, and a willingness to record the misses alongside the hits. This case was published with its resolution date stated up front for exactly that reason.
See the actual 3Dogs report
The full deliberated brief that produced the analysis above — the call, the probability assessment, the evidence classification, the panel’s position changes, and the preserved dissent. Case 2026-0026.
Open the full report (PDF) Start a Decision CaseQuestions this case answers
Why is a probability spread better than a single AI prediction?
A single number hides the uncertainty that should drive the decision. A spread tells you how much of the probability mass sits in each outcome, which is what you actually act on — and it can be scored later, so the forecaster is accountable.
What are calibrated AI predictions and Brier scores?
Calibration means that things you call 70% likely happen about 70% of the time. A Brier score measures that accuracy after the fact. We publish resolution dates specifically so our forecasts can be scored rather than quietly forgotten.
How does this compare with asking Gemini or ChatGPT to forecast?
We ran the identical question both ways and published both answers side by side. The single-model answer was fast, fluent and gave one number. The panel produced a spread, the reasoning behind each band, and preserved disagreement.
Try this on your own question.
Free, no card. Bring a real decision — ideally one where you already know the answer — and see what the panel does with it.
Start a decision casePeople also ask
What are calibrated AI predictions and Brier scores?
Calibration means things you call 70% likely happen about 70% of the time. A Brier score measures that accuracy after the fact. We publish resolution dates specifically so forecasts can be scored rather than quietly forgotten.
Related decision case studies
- AI Fraud Detection: Our Own Model Voted to Reject Us
Run on the MIT AI research scandal, our permanent adversarial seat voted against the panel's own premise - the clearest evidence the ch
- Can AI Analyze a Decision in Any Language?
A real EUR12B decision brief submitted in Greek and Mandarin, clarified in Pitjantjatjara and answered in French - the recommendation c
- Will Oracle Get a Bailout Over Its OpenAI Bet? The Odds
32 AI models across 12 vendor families and 3 clouds examined Oracle's BBB- downgrade and OpenAI concentration, and produced explicit od
- Pay the Ransom or Rebuild? 23 AI Models Split
A critical-infrastructure ransomware decision run across 23 AI models on three clouds. The panel split 10-to-9 and we published the spl