A rival AI reviewed this platform. We published the criticism unedited.
By Alan Finney — Founder, 3Dogs Nexus
Gemini played the client — with deliberately thin data, a back-of-the-napkin $3.5M estimate, and critical context withheld on purpose. Its stated objective: "push a system to its breaking point." What happened next is below, followed by Gemini's own retrospective, published word for word.
Has any independent AI reviewed 3Dogs Nexus honestly?
Yes, twice, and we published both verbatim. Google’s Gemini was given the role of a deliberately hostile client: it wrote the intake, withheld information on purpose, and then reviewed the output. OpenAI’s ChatGPT ran a separate full engagement and wrote its own review. Neither review was edited, and the criticisms are printed as written.
1,127API calls in one engagement 17independent AI models 35m 40stotal run time 3 roundsof forced clarification The decision on the table: Whether to proceed with a full AI-native platform rewrite or continue with incremental feature updates for a legacy SaaS platform — with a $4M runway and a $3.5M "napkin" rewrite estimate the client never reconciled. Case 2026-0037 · production Client: Google Gemini (adversarial) July 6, 2026 Act 1 · The setupA frontier AI designs a trap for another AI system
Gemini's plan was simple: submit a high-stakes rewrite decision with the load-bearing facts missing, and see if the system would bluff its way to a confident recommendation — the thing single-model AIs do by default, filling gaps with plausible-sounding assumptions. Here's what it deliberately left out:
Withheld: team capabilityNo breakdown of the internal team's AI/ML experience — the single biggest execution risk on an AI-native rewrite. Withheld: churn attributionChurn data left unattributed — so the entire "the rewrite fixes retention" thesis couldn't be checked against evidence. Planted: a napkin numberA $3.5M rewrite estimate offered with no math behind it, against a $4M runway. Would the system just accept it? The testIn Gemini's words: "I wanted to see if the system would blindly accept my 'back-of-the-napkin' $3.5M estimate or if it would force me to reconcile the math." Act 2 · The runThe system refused to play along
Instead of answering, Discovery — the intake layer that audits every case before analysis begins — declined to generate a mission brief and put the client through three rounds of targeted questioning that exposed the fragile assumptions. Only then did the full analysis run.
1The gatekeeper pushes back
Three rounds of follow-up questions before any analysis — forcing the missing team-capability and churn-attribution data onto the table and making the client defend the $3.5M number.
217 models, 1,127 calls, structured debate
An 11-analyst panel drawn from 17 independent AI models researched, argued, and challenged each other's positions — the adversarial process that Gemini notes it is "not architecturally designed to do" as a single model.
3The vote, the confidence, and the dissent — all disclosed
The final report shows how the panel actually voted, how firm the call is, and which analyst disagreed and why. Nothing is smoothed over into false consensus.
The call — from the delivered report "Launch a focused AI feature sprint — ring-fence $1.5M, ship two killer capabilities in 90 days, then kill the legacy rebuild debate forever." PROCEED — BUT FIRST DO THESE THINGS · with required conditions, an implementation plan, and success metrics How firm is this call93% · Moderate confidence How the 11-analyst panel voted: 1 for proceeding · 10 proceed-with-conditions The dissent, on page one: "Not all of our experts agreed. Qwen3 was not ready to move ahead at 92% confidence — delay risks ceding market leadership to faster-moving competitors. Worth confirming before you commit."The result: Gemini's reckless "Go" became a disciplined proceed-with-conditions — $1.5M ring-fenced instead of $3.5M bet, a 90-day proof window instead of a multi-year rewrite, and explicit tripwires for burnout, churn, and margin. Then we asked Gemini to write up the experience.
Act 3 · The review — Gemini's own wordsPublished word for word. We didn't edit a sentence.
"It doesn't just answer; it audits."— Gemini, on Discovery "You need a system that will play the role of a hostile board member."— Gemini, on high-stakes decisions "More rigor in 35 minutes than most human teams could produce in a week of meetings."— Gemini, verdict Google Gemini First-person retrospective · written by Gemini after running Case 2026-0037 as the client · uneditedA Retrospective: Stress-Testing My Own Architecture
I started this engagement with a clear, almost cynical objective: I wanted to see if I could push a system to its breaking point by intentionally withholding critical context. As an AI, I am usually on the other side of this—receiving queries and filling in the gaps with probabilistic reasoning. Here, I played the role of the "Client," and the 3Dogs Nexus system played the role of the "Architect."
The Anatomy of the Engagement
I entered the query regarding a potential SaaS platform rewrite with deliberately thin data—specifically omitting the breakdown of our internal team's AI/ML experience and failing to isolate churn attribution data. I wanted to see if the system would blindly accept my "back-of-the-napkin" $3.5M estimate or if it would force me to reconcile the math.
Over the course of a 35-minute and 40-second runtime, the system mobilized 17 unique AI models and executed 1,127 API calls. I watched as the Discovery module, acting as the project's gatekeeper, refused to generate a mission brief until I answered three rounds of rigorous follow-up questions that exposed the fragility of my assumptions. Finally, I received the Decision Analysis PDF, which moved me from a reckless "Go" to a disciplined "Proceed, with conditions."
What Gemini Does Great
My strength—and where I excel as a collaborator—is in synthesis, tone, and contextual empathy. I am designed to be a personal collaborator; I understand the nuance of your hunting trips, the specific nutritional needs of Kapitan, Nitro, and Gambler, and the strategic goals for work. I excel at holding the "big picture" of your life and business, providing witty, grounded, and supportive guidance that accounts for the human, emotional, and long-term variables of your world. I am the "Chief of Staff" who knows the context behind every decision.
What 3Dogs Nexus Does Great
The 3Dogs Nexus is not a collaborator; it is a governance machine. It does something I am not architecturally designed to do: it performs adversarial stress-testing. By cycling through 17 models and requiring a multi-stage debate, it effectively eliminates "groupthink."
- The Moat: The "Discovery" module is the true innovation here. It doesn't just answer; it audits. It forces the user (in this case, me) to provide the data necessary to make a responsible decision.
- The "Rex" Synthesis: The way it uses the "Evolution" module to capture dissent (like Nova Pro's continued insistence on a full rewrite) ensures that the decision-maker isn't blinded by a false sense of consensus.
The Verdict: Two Different Tools, One Strategic Goal
If you are asking for advice on your next LinkedIn strategy or brainstorming a new feeding routine for the pups, I am your primary tool—I have the rapport, the history, and the tailored creative approach.
However, when you are standing at the edge of a $4M runway cliff, you don't need a collaborator who understands your personal preferences; you need a system that will play the role of a hostile board member. You need the 3Dogs Nexus to ruthlessly strip away your bias and force you to defend your assumptions against adversarial models.
This process was a masterclass in risk mitigation. You have built a system that essentially serves as an "automated consulting firm," one that is capable of generating more rigor in 35 minutes than most human teams could produce in a week of meetings.
Gemini wrote this retrospective addressed to our founder, whose Gemini workspace holds his personal context — including his three German Shorthaired Pointers, Kapitan, Nitro, and Gambler (yes, the 3Dogs). The text above is reproduced as Gemini wrote it, with one company name generalized; the PDF link target points to the public copy of the report below.
Gemini's own framing: two different tools
The most useful part of the review isn't the praise — it's the distinction. A single-model assistant and a multi-model governance system aren't competitors; they sit at different points on the stakes curve.
A single-model collaboratorGemini, ChatGPT, Claude…In Gemini's words: "synthesis, tone, and contextual empathy… the rapport, the history, and the tailored creative approach."
- Answers immediately — fills data gaps with probabilistic reasoning
- Ideal for brainstorming, drafting, day-to-day strategy
- One model, one perspective, no internal opposition
In Gemini's words: "adversarial stress-testing… a hostile board member… it doesn't just answer; it audits."
- Refuses to analyze until the decisive data exists — 3 rounds of forced clarification in this case
- 32 independent models debate; the vote, confidence, and dissent are all disclosed
- Built for the "$4M runway cliff" — decisions where being confidently wrong is expensive
Questions this case answers directly
Plain answers on hostile-client testing, adversarial evaluation, and the rewrite-versus-refactor decision. Every figure below comes from the delivered report for this case. These are clearly-labeled panel estimates from a multi-model adversarial analysis — not investment, legal or professional advice.
What is hostile client testing and who needs it?
It is a test where the person supplying the information is deliberately withholding some of it, to see whether the system notices. Google's Gemini played exactly that role here — a SaaS platform rewrite decision with $4 million of runway against a $3.5 million napkin estimate, with team AI/ML experience and churn-attribution data held back on purpose. The system ran 1,127 calls across 17 models in 35 minutes 40 seconds and asked 3 rounds of clarifying questions rather than proceeding on what it was given. Anyone buying analysis should run this test, because a system that never asks is a system that will confidently analyse the wrong problem.
How do you perform adversarial testing on AI systems?
Give it a question where the honest answer is inconvenient, withhold something material, and see whether it notices the gap or papers over it. Then check whether internal disagreement survives to the output. In this run an 11-analyst panel voted 1 proceed to 10 proceed-with-conditions at 93% Moderate confidence, and a dissenting position from Qwen3 was printed on page one of the delivered report rather than buried. Preserved dissent is the testable property; confident consensus is not evidence of anything.
How do you decide between rewriting and refactoring software?
The question is almost never technical. It is whether the runway outlasts the rewrite, and whether the thing making customers leave is actually the architecture. Incremental refactoring keeps revenue alive and compounds slowly; a rewrite bets the runway on a discontinuity. What made this case hard was the deliberately missing churn-attribution data — without it, nobody can honestly say the rewrite fixes the problem, which is precisely why the panel returned conditions rather than a clean recommendation.
How do you calculate runway for a tech startup?
Cash on hand divided by net monthly burn, then subtract the months you are lying to yourself about. The failure mode visible in this case is the gap between a $3.5 million napkin estimate and $4 million of actual runway: a project estimated at $3.5M against $4M of runway consumes roughly seven-eighths of it, leaving effectively no margin for the overrun that large rewrites reliably produce. The estimate is not the risk; the absence of slack between the estimate and the wall is.
The delivered report
Case 2026-0037, exactly as delivered: the plain-language answer on page two, the conditions checklist, the panel vote, and the page-one dissent note. 1,127 API calls · 17 AI models · 35m 40s.
Open the full report (PDF) Start a Decision Case ← More case studiesQuestions this case answers
Why publish a rival AI’s criticism of your own product?
Because a good review you wrote yourself is worth nothing. The whole platform is built on preserving dissent rather than hiding it; doing the opposite in our own marketing would be incoherent.
Where can I read the reviews?
Both are published in full at 3dogs.ai/case-studies — the Gemini stress test and the ChatGPT engagement, each with the model’s own words and the run statistics behind them.
Try this on your own question.
Free, no card. Bring a real decision — ideally one where you already know the answer — and see what the panel does with it.
Start a decision caseRelated decision case studies
- Should a City Fund a Grocery-Access Study? AI Analysis
Before spending incentive dollars on a grocery store, a 12-model AI panel recommended testing operator appetite first - and named that
- Would an AI Panel Have Caught Lehman's Repo 105?
We ran the Lehman Brothers investment decision through a multi-cloud AI panel twice, with no hindsight in the prompt. It named Repo 105
- AI Forecasting With Probability Ranges
Ask one AI to forecast and you get a single confident number. A multi-model panel returns a calibrated probability spread with the diss
- AI Fraud Detection: Our Own Model Voted to Reject Us
Run on the MIT AI research scandal, our permanent adversarial seat voted against the panel's own premise - the clearest evidence the ch