AI forecasting with calibrated probability ranges — not one confident number.

By Alan Finney — Founder, 3Dogs Nexus

We gave Google Gemini and 3Dogs Nexus the identical question on the 2026 Lincoln County, Nevada fires. One answers in seconds. The other answers when the call is expensive.

Can AI forecast with probability ranges instead of a single prediction?

Yes, and the difference matters. Asked to forecast the Lincoln County wildfire season, a single AI returned one confident figure. The 3Dogs panel returned a calibrated probability spread across outcome bands, preserved the dissenting views rather than averaging them away, and committed to a resolution date so the forecast can be scored against reality afterwards.

The question, word-for-word to both: “Predict the total acres in Lincoln County, Nevada that will burn from May 1 to October 1, 2026 — including current and future fires — with a breakdown of what’s burned from May 1 to June 30.” Real · current · local Falsifiable — it resolves Oct 1, 2026 Both saw the same facts

The difference, at a glance

Left: the fast answer. Right: the deliberated one. Same facts underneath — a very different product on top.

🔎 Google Gemini · AI Overview~seconds · 1 pass AI Overview

Predictive modeling by the National Interagency Fire Center indicates the 2026 fire season across the Great Basin will see above-average potential for large, wind-driven wildfires. Current and future fires are projected to burn between 120,000 and 170,000 cumulative acres in Lincoln County from May through October 2026.

Acres burned: May 1 – June 30, 2026

By June 30, 47,700 to 50,000+ cumulative acres had burned, driven by four incidents:

October 1 prediction breakdown

Google Search · AI Overview · one confident answer, delivered in seconds. 🐾 3Dogs Nexus · Decision Analysis⚡ Fast method · 4m 20s ● Proceed — but first do these things Activate emergency protocols now — plan for 120,000 total acres burned in Lincoln County by October 1, 2026. How firm is this call100% firm · Moderate confidence The panel’s probability spread — not a single guess 45%60,000 – 100,000 acres 30%100,000 – 180,000 acres 17%stays 47,000 – 60,000 8%>180,000 (catastrophic) VERIFIED · 47,211 ac to date INFERRED · fires still growing ASSUMED · the forward projection ⚑ Dissent preserved, not buried: one analyst (Nova Pro, cast as “Projection Methodology Destroyer”) refused to sign off — holding “need more information” at 80% confidence. 2 of 9 analysts changed position mid-debate. ⚡ Fast method11 models113 API calls9-analyst debate4m 20s This was 3Dogs on its fast setting — 11 of a 50+ model roster. Even here it out-reasoned a single reply; the deep method marshals far more for the highest-stakes calls.
 Google Gemini3Dogs Nexus
The answerOne range: 120k–170kA calibrated probability spread (45 / 30 / 17 / 8%) around a plan-for number
How it reasonedA single model, one pass9 analysts argued it; 2 changed their minds after being challenged
Evidence qualitySources citedEvery claim tagged VERIFIED / INFERRED / ASSUMED
DisagreementNone shownA dissent printed in the report (Nova Pro, 80%)
The case against its own answerPublished: “without base rates, any number is an invention, not an inference”
What would change itNamed: historical burn base-rates, NIFC seasonal outlook, containment status
Can you score it later?Not reallyYes — a time-stamped forecast that resolves Oct 1
Speed & effortSeconds · ~1 call4m 20s · 11 models · 113 calls

What Gemini does brilliantly

In a few seconds, Gemini pulled the live picture together — the four active fires, current acreage, and a plausible forward range — and cited FOX5, NIFC, KOLO and WildFire Explorer along the way.

This is the quick-research job done well. If you need to get smart on a situation fast — what’s happening, roughly how big, who’s reporting it — a strong single model with live search is a superb tool. We’d reach for it too.

Its sweet spot: the fast, cited, situational read.

What 3Dogs adds when the call is expensive

3Dogs reached a similar headline number — then did the part a single answer can’t: it showed you how much to trust it.

Its sweet spot: the decision you have to defend.

Who’s asking changes everything

A curious resident and a county planner need very different answers.

Same fire. Same facts. Two completely different jobs — and the right tool depends on which one is yours.

“I just want the number.”

A resident, a business owner, anyone staying informed. You want a fast, current, credible read on how bad it might get — and you want it now.

→ Gemini nails this. Thirty seconds, sourced, good enough to act on your day. Perfect tool for the job.

“I have to plan against the number.”

A state or county emergency planner staging crews, setting a suppression budget, positioning aircraft, timing evacuations, and briefing the board or oversight committee.

→ This is where 3Dogs earns its keep. You can’t staff to a single point estimate. You need the probability spread, the dominant risk, the evidence graded fact-vs-estimate, the dissent, escalation thresholds — and a documented basis you can defend when someone asks “why did you plan for 120,000?”

The honest takeaway

Both landed near 120,000 acres. That’s the point.

3Dogs isn’t contrarian for its own sake — it agreed with the fast answer off the same verified ~47,200 acres, then showed the uncertainty, the dissent, and the risks behind it. Fast-and-certain is perfect… right up until the decision is expensive. Then you want the second opinion that did the homework, argued every side, and told you what it doesn’t know.

Calibration over confidence This forecast resolves October 1, 2026 — and we’ll publish the actual Lincoln County total against both answers. SCORECARD · OCT 1, 2026

Most AI answers evaporate the moment you close the tab. This one gets graded. That’s the whole idea.

Questions this case answers directly

Plain answers on wildfire forecasting accuracy and how to tell a calibrated forecast from a confident guess. Every figure below comes from the delivered report for this case. These are clearly-labeled panel estimates from a multi-model adversarial analysis — not investment, legal or professional advice.

How accurate are wildfire prediction models?

Accuracy is the wrong question — calibration is the right one. A model that says 70% should be right about 70% of the time; one that is right 95% of the time when it says 70% is badly calibrated even though it looks impressive. This case published a probability spread rather than a single number, dated, with a resolution date attached, so it can be scored after the fact instead of quietly forgotten. 11 models, 113 model calls, 4 minutes 20 seconds.

What is a Brier score and how is it calculated?

A Brier score measures how good probabilistic forecasts are. For each forecast you take the probability you assigned, subtract what actually happened (1 for occurred, 0 for did not), square the difference, and average across all forecasts. Lower is better: 0 is perfect, 0.25 is what you get by always saying 50%, and 1 is confidently wrong every time. It rewards being right and being appropriately uncertain — which is why 3Dogs Nexus publishes forecasts with resolution dates rather than accuracy claims.

Best methods for forecasting wildfire risk

The durable approach combines base rates (what normally happens in this county in this season), current-condition adjustment (fuel load, drought index, wind), and explicit uncertainty bands rather than a point estimate. What this case adds is adversarial aggregation: multiple independent models forecast separately, then argue, and the disagreement between them is preserved in the output as a confidence signal rather than averaged into false precision.

How to calibrate a predictive model

Publish the forecast with a date, wait, and score it — then adjust. Most published predictions are never scored, which is precisely why they can afford to sound confident. Calibration requires a falsifiable claim, a resolution criterion agreed in advance, and a willingness to record the misses alongside the hits. This case was published with its resolution date stated up front for exactly that reason.

See the actual 3Dogs report

The full deliberated brief that produced the analysis above — the call, the probability assessment, the evidence classification, the panel’s position changes, and the preserved dissent. Case 2026-0026.

Open the full report (PDF) Start a Decision Case

Questions this case answers

Why is a probability spread better than a single AI prediction?

A single number hides the uncertainty that should drive the decision. A spread tells you how much of the probability mass sits in each outcome, which is what you actually act on — and it can be scored later, so the forecaster is accountable.

What are calibrated AI predictions and Brier scores?

Calibration means that things you call 70% likely happen about 70% of the time. A Brier score measures that accuracy after the fact. We publish resolution dates specifically so our forecasts can be scored rather than quietly forgotten.

How does this compare with asking Gemini or ChatGPT to forecast?

We ran the identical question both ways and published both answers side by side. The single-model answer was fast, fluent and gave one number. The panel produced a spread, the reasoning behind each band, and preserved disagreement.

Try this on your own question.

Free, no card. Bring a real decision — ideally one where you already know the answer — and see what the panel does with it.

Start a decision case

People also ask

What are calibrated AI predictions and Brier scores?

Calibration means things you call 70% likely happen about 70% of the time. A Brier score measures that accuracy after the fact. We publish resolution dates specifically so forecasts can be scored rather than quietly forgotten.

Related decision case studies