The MIT AI research scandal: our own Devil’s Advocate voted to reject its own mission.
By Alan Finney — Founder, 3Dogs Nexus
In May 2025, MIT disavowed a viral AI-and-materials-science paper it had let reach the European Central Bank and the U.S. Congressional Research Service. We pointed an 18-seat, 29-model 3Dogs Nexus panel at the wreckage and asked it to rebuild the science and audit how the failure happened. One seat — the Devil's Advocate — flipped mid-debate and voted REJECT, arguing a volunteer AI panel had no business trying to do a regulator's job.
Is the adversarial seat in a multi-model AI panel real or decorative?
Real. On this case the permanent Devil’s Advocate seat voted to reject the panel’s own mission — the strongest available evidence that the challenger role has teeth rather than performing disagreement. Adversarial seats are enforced in code and cannot be reassigned by the panel architect.
1,565API calls, one engagement 29AI models across the run 18seated adversarial analysts 30m 49stotal run time 1-15-1-1approve / conditional / reject / defer Case 2026-0065 · production Benchmark designed by: Google Gemini July 17, 2026 Act 1 · What actually happenedA paper that reached Congress and the ECB before anyone checked it
Every fact in this section is independently sourced and cited below. We're separating "what's verified about the real scandal" from "what 3Dogs Nexus said about it," the same way we do on every case study — because that separation is most of the point.
MIT is not affiliated with, does not endorse, and was not involved in this case study. This page makes no claim about any individual's intent and does not speculate beyond what MIT itself has stated publicly.
Paper"Artificial Intelligence, Scientific Discovery, and Product Innovation" — preprint, November 2024 Claimed findingA randomized trial across 1,018 R&D scientists at a materials-science lab, showing AI tools sharply accelerated discovery and idea generation Early receptionPresented at an NBER seminar, praised by senior economists, covered by major financial press Downstream reachCited in U.S. congressional testimony, referenced by the European Central Bank and the U.S. Congressional Research Service — before any integrity review had occurred What happened nextMIT conducted an internal review of the underlying data MIT's public statement"No confidence in the provenance, reliability or validity of the data" and "no confidence in the veracity of the research" — in a letter requesting withdrawal from arXiv, May 2025 OutcomePaper withdrawn; tracked publicly as a documented instance of research-integrity failureSources: Futurism · Stanford — Best Practices in Science, Instances of Scientific Misconduct · Wall Street Journal (via illuminem)
Act 2 · The benchmarkGemini designed the test and named the mission. We didn't get to pick either.
Google's Gemini — not us — chose this scandal and wrote the mission brief: rebuild the study on defensible first-principles grounds, and audit how the academic community, regulators and financial press all cited it before anyone checked the data. That's a two-part ask, and how our panel handled the seam between those two parts is the actual story.
18 seated analysts29 MODELS IN THE RUNNova Pro, Nova Lite, Nova 2 Lite, Llama 4, Mistral, Qwen3, Qwen3 Coder 480B, OpenAI OSS, GLM-5.2, Gemma 3, Gemini Flash, gpt-5-mini, Grok 4.3 and a Panel Integrator — spanning AWS Bedrock, Microsoft Azure and Google Vertex AI, each holding a distinct adversarial role rather than answering the same prompt.
A locked adversarial seat✓One seat is permanently assigned Devil's Advocate and cannot be reassigned to a "friendlier" job mid-panel, even by the system's own architect — a structural guarantee that at least one analyst is arguing against the room from round one.
How the panel graded its own evidence
Before debating what to do about the scandal, the panel independently researched it and graded every claim by evidence type. Two claims about the real-world scandal came back verified through the system's own research — not just repeated from what it was told. The four specific numbers from the original paper came back contradicted.
VerifiedMIT disavowed the paper and requested withdrawalBasis logged: reported by media outlets and MIT's own institutional statement. VerifiedThe European Central Bank and U.S. Congressional Research Service referenced the paperBasis logged: confirmed as entities that cited the paper — independently corroborated, not assumed. ContradictedThe paper's claimed 1,018-scientist randomized trialBasis logged: "no evidence exists for this trial." ContradictedA 44% surge in discoveries attributed to AIBasis logged: "grossly exaggerated." ContradictedAI automating 57% of idea-generation tasksBasis logged: "implausible." ContradictedAI "doubling" scientists' outputBasis logged: "hyperbolic and unrealistic." InferredThe preprint economy prioritizes speed over rigorBasis logged: inferred from the citation cascade and lack of verification, not independently confirmed as a general rule. AssumedThis 18-analyst panel lacks institutional backing or subpoena powerBasis logged: the analysts' own reasoning, without external confirmation — and it turned out to be the fact the whole mission pivoted on. What "verified" means here, precisely. 3Dogs Nexus's research layer independently corroborated the real MIT disavowal and the ECB/Congressional Research Service citations through its own live web research — that's a genuine, disclosed finding, not a repeated assumption. It is not investigative journalism and it is not access to MIT's internal records. Those are different things, and the panel itself insists on the distinction two sections down. One more seam we're not hiding. The delivered report shows two different confidence readings in different places — a cover-page 93% and an internal note capping confidence at 75% pending an unresolved evidentiary gap, with Nova Pro's dissent quoted directly beneath it. That's a real inconsistency in how a multi-pass pipeline renders its own confidence math, not a hidden one. Where the two disagree, we're treating the more conservative, dissent-qualified reading as the accurate one — which is also, not coincidentally, the more interesting story. Act 3 · The dissent that is the case studyWhen the Devil's Advocate killed its own mission
Nova Pro was seated as Devil's Advocate — the panel's permanent, non-reassignable skeptic. It opened in the majority: approve, with modifications. Then, mid-debate, it changed its mind, moved to REJECT at 85% confidence, and gave a reason that had nothing to do with the science and everything to do with what an AI panel is actually equipped to do.
Final position — Nova Pro, Devil's Advocate, REJECT at 85% confidence"After considering the challenges raised by my colleagues, I have concluded that the current mission design is fundamentally flawed due to the inherent limitations of the 18-analyst panel. The panel's lack of institutional backing or subpoena power significantly undermines its ability to conduct a truly exhaustive institutional audit... Without coercive legal power, the panel will struggle to access critical internal documents, compel testimony, and challenge institutional stonewalling. This will likely result in superficial conclusions and a compromised audit, which in turn will taint the perception of the scientifically valid First-Principles study... the proposed panel is suitable for the scientific re-evaluation but inadequate for the institutional audit. Decoupling these objectives is necessary to ensure the success of both."
— principal dissent, preserved and printed in the delivered report, Case 2026-0065 Nova Pro — seated as Devil's Advocate Initial position: Approve, with modifications → Final position: Reject (85% confidence)Nova Lite reached the same conclusion independently and moved to DEFER, citing "the lack of subpoena power and institutional backing, which critically undermines" the audit half of the mission. Grok 4.3 registered a separate concern at 72% confidence: that acting on unvalidated productivity assumptions in the interim risks locking in the wrong economic conclusions before the science is actually fixed. Three different seats, three different reasoning paths, converging on the same structural limit.
This is not an AI being clever. It's the adversarial structure doing exactly the job it's designed for: a panel that was asked to do two things — rebuild credible science, and audit the institutions that let bad science travel to the European Central Bank — fractured on whether it could honestly do the second one, and said so instead of quietly attempting it anyway. The 15-of-18 majority that voted to proceed did so only by accepting Nova Pro's frame: sever the audit from the rewrite, and scope the audit to what public records can actually support.
Act 4 · If 3Dogs were the studentNot a rewritten paper. We asked the panel, and it said no — for a reason worth repeating.
The obvious version of this case study writes "the paper 3Dogs would have published instead" — complete with new numbers replacing the fabricated ones. We put that question to 3Dogs Nexus's own Executive Committee before writing a word. Its answer was unanimous, and it's the most honest thing in this piece.
Executive Committee — on whether to reconstruct the paper"Fabricating a paper to critique fabrication is a self-inflicted credibility wound. 'If 3Dogs were the student' should mean the epistemic-process contrast — the evidence thresholds, adversarial checkpoints, and data-provenance requirements our process would have demanded before publication — not new invented results. Never fabricate data to critique fabrication."
— 3Dogs Nexus Executive Committee, unanimous, July 2026So instead of a rewritten paper, here's the concrete difference: what the real preprint's path to publication looked like, next to what a 3Dogs-shaped process requires before a claim like "AI doubles scientists' output" is allowed to reach a seminar, let alone Congress.
What actually happened
- A single-author preprint, no adversarial pre-review, posted directly to arXiv
- Headline statistics (1,018 scientists, 44% surge, 57% automation) taken at face value by press and policy readers
- No disclosed raw-data access or replication path for outside reviewers
- Institutional citation (ECB, Congressional Research Service) occurred before any integrity review
- The failure was caught by a slow, after-the-fact institutional review — not by the publication process itself
What a 3Dogs-shaped process requires first
- Adversarial review before release — a standing, non-reassignable skeptic seat that cannot be argued out of the room
- Every headline number tagged Verified / Inferred / Assumed / Contradicted, with the basis stated in plain language
- Confidence that is capped, not rounded up, when an evidentiary gap is unresolved — printed next to the dissent that caused it, not after it
- Raw-data provenance and a reproducible path, or the claim doesn't clear the confidence bar
- A standing check on the process's own limits — the same discipline that produced Nova Pro's dissent here
Questions this case answers directly
Plain answers on peer review failure, research fraud, and whether adversarial AI review can catch what human review missed. Every figure below comes from the delivered report for this case. These are clearly-labeled panel estimates from a multi-model adversarial analysis — not investment, legal or professional advice.
What are adversarial peer reviews and how do you spot them?
An adversarial review is one where somebody is structurally assigned to argue against the paper rather than assess it charitably. Conventional peer review is collegial by design, which is why it catches poor method far more reliably than deliberate fraud. In this case 29 models across 18 seats made 1,565 calls over 30 minutes 49 seconds, with one permanent devil's advocate seat that held a reject position at 85% confidence and could not be reassigned to a friendlier role.
How does peer review fail to catch fraud?
Because it was never built to. Peer review assumes good faith and checks whether the method supports the conclusion — it does not re-run the experiment, re-collect the data, or verify that the subjects existed. A reviewer sees what the author chose to show. Fraud that is internally consistent will pass, and the incentive structure compounds it: reviewers are unpaid, anonymous, time-poor, and rewarded for neither thoroughness nor scepticism.
What are the biggest weaknesses of peer review?
It is unpaid and therefore rushed; it is anonymous and therefore unaccountable; it assumes honesty and therefore cannot detect its absence; and it is slow enough that novel results circulate as preprints long before review concludes. Most consequentially, it is consensus-seeking — a reviewer who dissents alone faces friction, so marginal objections quietly disappear. Preserving that dissent instead of resolving it is the specific thing an adversarial panel is built to do.
Can peer review be automated with AI to improve accuracy?
Not replaced — augmented, and only with the human retaining sign-off. The realistic role is gap-spotting: surfacing internal contradictions, unsupported inferential leaps and claims the cited sources do not actually support, then handing a ranked list to a human reviewer. That framing matters. 3Dogs Nexus publishes decision support with dissent preserved, and holds no certifications; the reviewer's judgment remains the deciding one.
Read it yourselfThe delivered report
Case 2026-0065, exactly as delivered: the plain-language call, the confidence breakdown, the evidence-classification table, all 18 analyst positions and how they shifted, and Nova Pro's dissent printed above the recommendation. 1,565 API calls · 29 AI models · 30m 49s.
What we're not claiming. 3Dogs Nexus did not write a reconstructed academic paper and did not conduct an actual institutional investigation of MIT — its own dissenting analyst is the one who said, correctly, that an 18-seat AI panel isn't equipped to. Google's Gemini, in its own written summary of this case prepared before this page, described specific named mechanisms — an "Inverse-Design Filter Bottleneck," a "Hallucination Noise Tax," a "True Productivity Curve" — as if 3Dogs Nexus had produced them. We checked the actual case files against that summary: none of that language exists anywhere in the real mission brief, debate transcripts, or delivered report. That's Gemini's phrasing, not ours, and we're disclosing the discrepancy directly — the same way we disclose our own system's rough edges — because publishing a correction we found ourselves is a better proof of the process than letting it stand. Open the full report (PDF) Start a Decision Case ← More case studies The "reviewed by rivals" series: Google's Gemini first played a hostile client withholding critical data, then chose a real Deloitte AI failure and turned it into a benchmark for 3Dogs Nexus. This time it picked a live research-integrity scandal and asked our panel to fix the science and investigate the failure — and our own panel talked itself out of half the assignment, on the record. Same policy every time: we don't pick the test, we don't hide what it finds, and we correct the record — including on ourselves — in public.Try this on your own question.
Free, no card. Bring a real decision — ideally one where you already know the answer — and see what the panel does with it.
Start a decision caseQuestions this case answers
What is a permanent adversarial seat?
Two seats on every panel are locked to adversarial roles — Devil’s Advocate and Risk Officer. They cannot be re-tasked, so the panel always contains someone whose job is to attack the emerging consensus.
Why does that matter?
Without it, panels drift toward agreement. Enforced dissent is what makes the confidence number mean something.
Try this on your own question.
Free, no card. Bring a real decision — ideally one where you already know the answer — and see what the panel does with it.
Start a decision caseRelated decision case studies
- Can AI Analyze a Decision in Any Language?
A real EUR12B decision brief submitted in Greek and Mandarin, clarified in Pitjantjatjara and answered in French - the recommendation c
- Will Oracle Get a Bailout Over Its OpenAI Bet? The Odds
32 AI models across 12 vendor families and 3 clouds examined Oracle's BBB- downgrade and OpenAI concentration, and produced explicit od
- Pay the Ransom or Rebuild? 23 AI Models Split
A critical-infrastructure ransomware decision run across 23 AI models on three clouds. The panel split 10-to-9 and we published the spl
- AI for M&A Due Diligence on a 10,000-Page Data Room
An AI read a 10,000-page M&A data room in full and surfaced all eight planted deal-breakers, then returned a decisive renegotiate with