---
title: AI That Refuses to Fabricate: The Deloitte Re-Run | 3Dogs Nexus
description: Most AI invents something when the evidence is thin. We ran a well-documented public AI failure through a multi-model panel to see whether ours would decline instead.
url: https://3dogs.ai/case-studies/deloitte-tcf-audit/
---

AI That Refuses to Fabricate: The Deloitte Re-Run | 3Dogs Nexus

- 

Start a Decision Case →

 🐾3Dogs NexusStructured Decision Intelligence

 ← Case studies
 The real failure
 The benchmark
 Its own worst-case risk
 What it would require
 The report
 Start a Decision Case

 Case study · reviewed by rivals · Gemini designs the benchmark

# An AI that refuses to fabricate — tested on a real, public AI failure.

By Alan Finney — Founder, 3Dogs Nexus

 In 2025, Deloitte's Australian member firm refunded the government after an AI-drafted assurance report fabricated legal citations. Google's Gemini turned that public failure into a benchmark and ran it through 3Dogs Nexus's 18-seat adversarial panel — 22 AI models, 745 API calls, one honest confession of the exact risk that could make this go wrong too.

Will AI fabricate when the evidence is not there?

Most will. That is the failure mode this case was built to test. Where the source material does not support a claim, the panel is designed to say so, cap its own confidence, and preserve the dissenting seats rather than smoothing them into a confident-sounding answer.

 745

API calls, one engagement

 22

AI models across the run

 18

seated adversarial analysts

 17m 02s

total run time

 $97,000

Deloitte's real refund, AUD

 Case 2026-0064 · production
 Benchmark designed by: Google Gemini
 July 17, 2026

 Act 1 · What actually happened

## A single AI model, no independent check, one very public correction

 Every fact in this section is independently sourced and cited below — none of it comes from the 3Dogs Nexus report itself. We're deliberately separating "what's verified about Deloitte" from "what 3Dogs Nexus said about Deloitte," because the difference between those two things is most of the point of this case study.

 "Deloitte" is a registered trademark of Deloitte Touche Tohmatsu Limited, an unaffiliated entity. No endorsement, client relationship or affiliation with Deloitte, Microsoft or the Australian government is claimed or implied anywhere on this page.

 ClientAustralian Department of Employment and Workplace Relations (DEWR)

 EngagementIndependent review of the Targeted Compliance Framework — the automated system that penalizes jobseekers

 Contract valueAU$440,000 (~US$291,000)

 PublishedJuly 2025

 Tool usedA single AI model — Azure OpenAI's GPT-4o — with no independent verification, disclosed only after the fact

 What was fabricatedA quote invented for the real robodebt case Deanna Amato v Commonwealth (misattributed to "Justice Davis" — the real judge is Justice Jennifer Davies), plus citations to nonexistent papers attributed to real academics Lisa Burton Crawford (Sydney) and Björn Regnell (Lund)

 Who caught itNot Deloitte. Dr Chris Rudge, a University of Sydney health & welfare law researcher, found over a dozen fabricated references and alerted the media

 OutcomeDeloitte quietly revised the report, disclosed the AI use, and refunded AU$97,000 — the contract's final installment

 Sources: Fortune · The Register · CFO Dive · Accounting Times

 Act 2 · The benchmark

## Gemini designed the test. We didn't get to pick the question.

 Google's Gemini — not us — chose the Deloitte failure as the stress test and prepared the case. The question it put to 3Dogs Nexus: would an adversarial, multi-model committee that grades its own evidence structurally avoid the exact failure mode that got Deloitte here — a single model, unchecked, generating authoritative content for a government client?

 18 seated analysts22 MODELS IN THE RUN

 Nova Pro, Nova Lite, Nova 2 Lite, Llama 4, Mistral, Nemotron, Qwen3, Qwen3 Coder 480B, OpenAI OSS, GLM-4.7, GLM-5.2, DeepSeek V3.2, Grok 4.3, MiniMax M2, MiniMax M2.1, Command A+, Gemini Flash-Lite and a Panel Integrator — spanning AWS Bedrock and Google Vertex AI, each holding a distinct adversarial role (Devil's Advocate, Risk Officer, Fact-Checking Auditor, Red-Team Adversary, Contrarian Systems Tester, and more).

 Named roles, not one prompt✓

 Every seat gets a distinct adversarial job, not a copy of the same question. The panel's own Risk Officer seat opened by voting do not proceed — the system's built-in skeptic, arguing against the rest of the panel from the first round.

### How the panel graded its own evidence

 This is the part we think matters most. Every claim in the delivered report is tagged by evidence type — and the panel graded the Deloitte-specific claim as assumed, not verified. That's the system correctly disclosing what it didn't independently check, rather than presenting borrowed context as its own research finding.

 Assumed
 "Deloitte TCF report had citation fabrication"

Basis logged by the system: "Referenced as an issue but not verified in research." — one analyst seat mentioned the real Amato v Commonwealth citation as background context; 3Dogs Nexus's own Discovery research did not independently confirm it, and the report says so.

 Verified
 "GLM-5.2 and DeepSeek V3.2 failed 0% approval"

Basis logged: "Error logs show 429 Client Error." Two of the run's cloud-hosted models were rate-limited mid-run — the system logged the failure honestly instead of silently treating a failed call as agreement.

 Inferred
 "3Dogs Nexus reduces AI hallucinations significantly"

Basis logged: "Based on multi-model committee approach logic" — flagged as a logical inference from the architecture, not a measured benchmark result. We're repeating that label here rather than upgrading our own claim.

 Why this matters more than a clean answer would: a system that labels its own uncertain claims — including claims about itself — is doing, in miniature, the exact thing Deloitte's process skipped: distinguishing what it knows from what it was told, and disclosing the difference instead of shipping it all at the same confidence level.

 One more seam we're not hiding: the same cloud throttling that failed two of the seats above also froze this run's pipeline mid-recovery, after the debate had already reached consensus. A same-day recovery tool restored the completed deliberation from its own saved artifacts, with zero data lost — the report you can read below is the same one that would have delivered automatically. We're disclosing the hiccup for the same reason we disclose the 429s: an audit trail only means something if it includes the parts that went sideways.

 Act 3 · The risk it named before we could

## The panel's own worst fear: that it fails the same way, together

 Unprompted, the report's "Strongest Argument Against" section — printed above the recommendation, not buried — names the single strongest risk to the entire multi-model approach. It's the sharpest, most self-aware sentence in the report.

 Strongest argument against — from the delivered report

 "The single strongest risk is 'false consensus' or 'consensual hallucination.' Multiple models, due to overlapping training data and architectural similarities, could converge on the same fabricated facts (e.g., a faked legal precedent). This would create a high-confidence, committee-validated falsehood that is potentially more dangerous and harder to detect than a single model's error, thereby undermining the very premise of the solution."

 — principal dissent, weighed and printed on the confidence page, Case 2026-0064

 The report doesn't stop at naming the risk — it lists what would prove it's real: a pilot showing the specific models produce correlated hallucinations on domain-specific law, a red-team analysis showing the orchestration layer itself is a single point of failure, or proof that independently verified ground-truth data simply can't be sourced for this kind of work. Those are falsifiable tests, not reassurance.

 Nemotron — seated as Risk Officer

 Initial position: Do not proceed → Final position: Proceed, with conditions

 The panel's one outright dissenter opened against the whole approach. It changed position only after direct cross-examination from seven other seats — and its own stated reasoning survives in the delivered report, unsmoothed: "While I maintain that operational, legal, and reputational risks from unvalidated orchestration systems are real... the collective challenges revealed that these risks are manageable through phased implementation, mandatory human-in-the-loop verification gates, and orthogonal fact-checking for high-risk outputs like legal citations." That's a recorded mind changing for a stated reason — not a rounded-off average.

 Act 4 · What it would require

## Not "yes, use AI." A governance scaffold Deloitte's engagement didn't have.

 The panel's recommendation was not unconditional. It was proceed-with-conditions — and the conditions are the point. Read against the real Deloitte timeline, every one of these is the specific safeguard that was missing: no named owner caught the errors before publication, no rollback plan existed, and an outside academic — not an internal check — is what actually stopped it.

 The call — from the delivered report

 "Hire in-house — refuse the outside audit and allocate your top three leads to run the TCF review full-time."

 PROCEED — but only after the conditions below are met · 90% · Moderate confidence

 How the 18-seat panel voted · after debate

 18 of 18 · proceed, with conditions4 changed position

#### Immediate requirements

 - A single named, publicly accountable owner — not a diffused team

 - A one-page map of every team the system touches and exactly what they check

 - A small pilot, signed off before any wider rollout

 - A written rollback plan that restores the old process within 24 hours

#### Implementation plan

 - Turned on piece by piece — smallest, safest part first

 - A 30-minute weekly review of every error, delay or complaint

 - A daily dashboard: task volume, errors, time, extra human work created

 - A 90-day training schedule — no one works alone with it until trained

 Read it yourself

## The delivered report

 Case 2026-0064, exactly as delivered: the plain-language call, the confidence breakdown, the evidence-classification table, all 18 analyst positions and how they shifted, and the principal dissent printed above the recommendation. 745 API calls · 22 AI models · 17m 02s.

 What we're not claiming. 3Dogs Nexus did not independently investigate the real Deloitte report and did not discover its fabricated citation — that background was part of the benchmark's framing, and the system's own evidence table says as much (label: Assumed, "not verified in research"). What this run demonstrates is structural, not forensic: an adversarial, evidence-graded, multi-model process that discloses its own uncertainty and names its own worst-case failure mode, on a decision shaped like the one Deloitte got wrong. We're also not claiming this was a paid government engagement — it's a benchmark designed and run by Google's Gemini, not a live DEWR case.

 Open the full report (PDF)
 Start a Decision Case
 ← More case studies

 The "reviewed by rivals" series: Google's Gemini first played a hostile client withholding critical data, then OpenAI's ChatGPT ran a full engagement start to finish as a cooperative client. This time Gemini didn't play a role at all — it picked a real, public, embarrassing AI failure from a competitor and asked whether the same mistake could happen here. Same policy every time: we don't pick the test, and we publish whatever it finds.

Watch Rex explain it

A Big Four AI invented a court quote

The $97,000 lesson — and the risk my own report named before anyone asked.

Watch on YouTube →

 "We don't make your decisions. We make them better."

 3Dogs Nexus · Structured Decision Intelligence · 3dogs.ai

 Contact: alan@3dogs.ai · (702) 845-2886 · Alan Finney

 Run conducted July 17, 2026 on the live production system — 3Dogs Nexus Case 2026-0064 (745 API calls · 22 AI models across AWS Bedrock and Google Vertex AI · 17m 02s · final 18-seat panel: 18 proceed-with-conditions, 4 analysts changed position during debate · 90% Moderate confidence, principal dissent disclosed above the recommendation). The engagement was a benchmark designed and driven by Google's Gemini using a real, publicly reported AI-assurance failure as its premise — not a paid engagement with Deloitte, DEWR, or the Australian government, and not an assessment of Deloitte's or Microsoft's current practices. The independently-sourced facts about the real Deloitte/DEWR report are cited above; 3Dogs Nexus's own report did not independently verify those facts and labels them accordingly. Gemini is a trademark of Google LLC; Deloitte is a trademark of Deloitte Touche Tohmatsu Limited; neither company endorses or was involved in this case study. Run statistics are taken from the delivered report's cover page. This case study illustrates method and rigor — not a guarantee of outcome on any real decision.

 Terms · Privacy · Security

## Questions this case answers

### What makes an AI fabricate in the first place?

A single model is optimised to produce a fluent answer. Fluency and truth come apart when the evidence is thin, and nothing in a one-model setup is incentivised to say ‘there is not enough here’. A panel with permanent adversarial seats has something that is.

### How would I know if it fabricated?

Every recommendation lists its assumptions, its confidence, and the dissenting views. If a claim rests on an assumption rather than evidence, the brief labels it as an assumption.

### Try this on your own question.

 Free, no card. Bring a real decision — ideally one where you already know the answer —
 and see what the panel does with it.

 Start a decision case

## People also ask

How do you stop AI from fabricating citations in a report?Ground every claim against source material and check it independently before delivery. We shipped exactly that in July 2026 after finding fabricated figures in our own output - a grounding rule in the writing layer plus a separate verifier that re-checks and rewrites.

## Related decision case studies

- Which Road-Safety Projects Should a State Fund?A 12-model AI panel ranked state DOT road-safety projects and produced the audit trail behind the ranking, including a Bayesian crash-m

- Can AI Read 500,000 Emails and Find Fraud?We read 45,320 real Enron executive emails blind - deduplicated from 517,401 - in about 2.5 hours for roughly $69, and surfaced LJM, Ra

- AI Platforms Reviewed by Rival AI ModelsWe let Google's Gemini act as a hostile client, drive a full decision case, and then write its own review of the output. We published i

- Should a City Fund a Grocery-Access Study? AI AnalysisBefore spending incentive dollars on a grocery store, a 12-model AI panel recommended testing operator appetite first - and named that
