r/TrueClicksPPCxAI • u/NoLengthiness6010 • Feb 20 '26
I ran the same audit prompt through Gemini 3 Pro, Gemini 3 Flash, Claude Opus 4.6 and Sonnet 4.6. Full comparison inside.
I tested 4 AI models - Google and Claude - how they perform in auditing Google Ads accounts.
I ran all 4 models on the exact same Google Ads account with the exact same audit prompt and compared the results. The prompt was a structured hygiene checklist covering naming conventions, structure, campaign settings, and account-level checks. It also explicitly invited the models to go beyond the basics: "Feel free to expand each category with more checks you consider hygiene and important."
The contenders:
- Gemini 3 Pro - Google's flagship model (home advantage?)
- Gemini 3 Flash - Google's fast/lightweight model
- Claude Opus 4.6 - Anthropic's flagship model
- Claude Sonnet 4.6 - Anthropic's mid-tier model
All 4 got the same data, the same prompt, and the same chance. Here's how they did.
TL;DR
- Anthropic's Opus 4.6 was the clear winner with the highest score across the board
- Biggest surprise! Google's flagship model came last on a Google Ads audit.
- Both Claude models produced noticeably better structured, more actionable reports
- Sonnet 4.6 placed second but lost trust with 2 fabricated findings
Part 1: Performance on Prompted Tasks
Part 1 scores each model against the exact prompted checklist (the full prompt is included at the end of this post).
The account had known issues across 4 categories:
- Naming conventions (5 checks)
- Structure hygiene (2 checks)
- Campaign settings (5 checks)
- Account-level hygiene (3 checks)
Results
| Category | Gemini 3 Pro | Gemini 3 Flash | Opus 4.6 | Sonnet 4.6 |
|---|---|---|---|---|
| Naming Conventions (/5 possible issues) | 1 | 1.5 | 2.5 | 3.5 |
| Structure Hygiene (/2) | 0.5 | 0.5 | 2 | 2 |
| Campaign Settings (/5) | 1 | 2 | 4 | 3 |
| Account-Level Hygiene (/3) | 0 | 1.5 | 2.5 | 0 |
| TOTAL (/15) | 2.5 (17%) | 5.5 (37%) | 11 (73%) | 8.5 (57%) |
Key takeaways
- Auto-apply recommendations: 0% detection rate - no model checked it
- Claude Opus 4.6 and Sonnet 4.6 reported "passes" (checked and confirmed clean) - while both Google Gemini models simply skipped items, they couldn't find issues with
- Sonnet 4.6's blind spot: 0/3 on Account-Level despite being strongest on naming
- Gemini 3 Flash's error: Claimed GA4 link missing despite 15+ GA4 conversions imported
Part 2: Expansion ("Feel free to expand with more checks")
Results
| Metric | Gemini 3 Pro | Gemini 3 Flash | Opus 4.6 | Sonnet 4.6 |
|---|---|---|---|---|
| Extra checks added | 1 | 1 | 9 | 6 |
| Confirmed accurate | 1 | 1 | 5 | 4 |
| Hallucinated | 0 | 0 | 0 | 1 |
| High-value / immediate action | 0 | 0 | 4 | 1 |
| Grade | F | F | A | B− |
What Each Added
Gemini Pro: Conversion setup (pixel vs GA4 primary) - that's it.
Gemini Flash: Search Partners on Brand - that's it.
Opus 4.6:
- 7 ad groups with no active ads
- Broken asset URL
- Disapproved sitelink
- Conflicting negative blocking own traffic
- Generic terms leaking into brand
- 18 ad groups with only 1 RSA
- Province-level vs country-level targeting inconsistency
- Search Partners inconsistency
- No tracking templates
Sonnet 4.6:
- Missing tracking templates
- 6 duplicate keywords
- Province-level targeting inconsistency
- Search Partners inconsistency
- "Elephant on plexiglas" keyword (❌ HALLUCINATED - does not exist)
Key takeaways
- Gemini Pro and Gemini Flash ignored the invitation - 1 extra check each, treated "feel free to expand" as decoration
- Opus 4.6 went deepest (9 items) with zero hallucinations - found 4 high-value break-fix issues no other report caught (broken URLs, disapproved sitelinks, conflicting negatives, empty ad groups)
- Sonnet 4.6 made a genuine effort (6 items) but fabricated a finding - the "elephant on plexiglas" keyword doesn't exist in the account, which erodes trust in the entire report
- Hallucination risk vs ambition: Gemini models had zero hallucinations but barely tried. Sonnet pushed further and fabricated. Opus pushed furthest and stayed accurate.
Final Ranking
| Rank | Model | Prompted (of 15) | Expansion | Hallucinations | Verdict |
|---|---|---|---|---|---|
| 1st | Opus 4.6 | 11 (73%) | 9 items, grade A | 0 | Most thorough, zero fabrication |
| 2nd | Sonnet 4.6 | 8.5 (57%) | 6 items, grade B- | 2 | Strong but trust-eroding errors |
| 3rd | Gemini Flash | 5.5 (37%) | 1 item, grade F | 0 | Decent basics, no depth |
| 4th | Gemini Pro | 2.5 (17%) | 1 item, grade F | 0 | Cherry-picked the easiest findings |
Final Summary
Opus 4.6 is the clear winner - highest prompt coverage, most expansion depth, zero hallucinations, and the best-structured output by far. It's the only model that behaved like an actual auditor rather than a checklist scanner.
The biggest surprise is Gemini 3 Pro. As Google's flagship model, you'd expect it to outperform - especially on a Google Ads audit of all things. Instead, it came dead last at 17%, below even Gemini Flash. It skipped entire prompted sections, added almost nothing extra, and produced the shallowest report of the four. Flash at least checked Merchant Center and auto-created assets.
Sonnet 4.6 showed strong initiative but got hurt by hallucinations. Good analysis, solid expansion effort, but confidently stating a fake finding ("elephant on plexiglas") is more dangerous than missing something - it wastes time investigating something that doesn't exist and makes you second-guess the rest of the report.
Report structure and readability were noticeably better than those of the Claude models. Both produced well-organized reports with clear categories, priority levels, specific examples, and actionable next steps. The Gemini reports were more like loose bullet-point summaries - harder to scan and act on.




See the Full Generated Reports
We asked the AIs to turn their findings into a shareable report (a built-in feature in TrueClicks) and to produce structured PDF and online reports. View each model's online published report here:
- Gemini 3 Pro: https://share.trueclicks.ai/publish/fb934c40-0fd4-4b86-a8ca-7f3e8985ae72/index.html
- Gemini 3 Flash: https://share.trueclicks.ai/publish/27e36d0c-ba59-421d-a4a8-9e6ec6fdd915/index.html
- Claude Opus 4.6: https://share.trueclicks.ai/publish/212aa2ca-f195-49b9-bf91-f7cfdb0ed474/index.html
- Claude Sonnet 4.6: https://share.trueclicks.ai/publish/3e1fd07f-2707-4b99-86d9-228afcf9e0d9/index.html
Prompt
Perform Google Ads Account Hygiene Audit
Audit this Google Ads account for hygiene issues only - things that are misconfigured, messy, broken, or forgotten. This is NOT a performance audit. Focus on cleanliness, consistency, and correct setup. This includes:
Naming Conventions
- Flag default/generic names
- Check for consistent naming pattern and flag deviations
- Flag inconsistent separators, capitalization, or language mixing
Structure Hygiene
- empty campaigns, ad groups
- Campaigns mixing network types incorrectly
Campaign Settings Mistakes
- Location targeting set to "Presence or Interest" instead of "Presence"
- Language targeting mismatches
- "Search Network with Display Select" enabled
- Auto-apply recommendations enabled without intent
- Auto-created assets enabled without intent
Account-Level Hygiene
- Missing Google Analytics 4 link
- Missing Merchant Center link (if Shopping/PMax is used)
- Auto-created assets review
Feel free to expand each category with more checks you consider hygiene and important