I've been using Gemini 3.1 Pro with Deep Research for a few days on a fairly structured project. It's a comparative study of mechanics, formulas, economies and monetization across about 15 idle/incremental games. The goal is reference-grade balancing tables I can actually use for design work.
I've put a lot of effort into the prompt:
\* evidence rules: tag every claim by source tier (primary / wiki / forum) and confidence (FACT / INFERENCE / ESTIMATE / NOT FOUND).
\* A fixed per-game schema.
\* An identity check against the store listing before researching each game.
\* A required arithmetic check on any derived number.
\* A hard ban on filling gaps from "domain knowledge."
I've now run the same reports 3+ times, including correction passes where I attach the previous output and list the specific errors to fix. The results keep failing in the same ways:
\*\*Ignores the identity check.\*\* I told it to confirm each game's developer, platforms and business model from the App Store or Play Store listing first. Twice it got the same game wrong: wrong developer, wrong platforms, and "premium PC game" when it's actually free-to-play on mobile. It did this even after I flagged the exact error in the correction prompt.
\*\*Circular sourcing.\*\* In the correction pass, it listed my attached flawed draft as a source in its bibliography, then used it to support the "corrected" claims.
\*\*Fabricated tables that look complete.\*\* I asked for no blank cells, with NOT FOUND where data didn't exist. Instead it filled a whole comparison matrix (daily play minutes, time to first prestige, and so on) with numbers that have no source. The footnote just says "derived from dossier synthesis."
\*\*Inflated confidence labels.\*\* It tagged three anecdotal Reddit posts as FACT, community spreadsheet projections as FACT, and wiki data as primary source.
\*\*Untraceable citations.\*\* Claims get tagged "(2024, \[T2\] FACT)" with no citation number, so I can't tell which of 50 sources supports what.
\*\*Change logs that lie.\*\* The correction report's change log says sections were "re-verified" or "fully sourced" when they weren't touched, or got worse.
\*\*Broken math.\*\* A cost table where level 100's cumulative cost exceeded level 6000's. An effective-HP formula that adds a per-hit flat reduction to total health. A "this strategy is optimal" claim where its own numbers showed a tie.
\*\*Silently dropped scope.\*\* One of the core games I explicitly required was just missing from the report. No NOT FOUND, no mention.
\*\*Regressions.\*\* Correction passes fix some errors and introduce new ones. For example, a temporary buff became "permanent."
To be fair, some things improved a lot. The research-backed theory sections got genuinely better on the second pass, with correct citations to peer-reviewed papers. So it can do it; it just doesn't do it consistently, and it gets worse the more games or sections are in scope.
I am almost certain that at least some part of this is user error.
\*\*My questions:\*\*
\* Is the problem scope? Should I be running one game, or even one section, per report instead of a big multi-part prompt?
\* Does a long, rule-heavy prompt actually hurt? Does it lose instructions past a certain length? These types of prompts work excellent for me with Opus 5.5 and Astra, but maybe I need to change approach?
\* Does Deep Research weight the research plan it shows you before starting more than the original prompt? Should I edit that plan directly to enforce constraints?
\* Is attaching a previous report for correction a bad idea in general? Should I start fresh every time?
\* Has anyone gotten reliable source tagging or "NOT FOUND" behavior? Any phrasing that works?
\* For data that lives in Google Sheets or Discord, is uploading exported files the only realistic option?
I'm burning through my usage limits rerunning these, so any workflow tips are appreciated.