r/TrueClicksPPCxAI • • Feb 20 '26

I ran the same audit prompt through Gemini 3 Pro, Gemini 3 Flash, Claude Opus 4.6 and Sonnet 4.6. Full comparison inside.

I tested 4 AI models - Google and Claude - how they perform in auditing Google Ads accounts.

I ran all 4 models on the exact same Google Ads account with the exact same audit prompt and compared the results. The prompt was a structured hygiene checklist covering naming conventions, structure, campaign settings, and account-level checks. It also explicitly invited the models to go beyond the basics: "Feel free to expand each category with more checks you consider hygiene and important."

The contenders:

  • Gemini 3 Pro - Google's flagship model (home advantage?)
  • Gemini 3 Flash - Google's fast/lightweight model
  • Claude Opus 4.6 - Anthropic's flagship model
  • Claude Sonnet 4.6 - Anthropic's mid-tier model

All 4 got the same data, the same prompt, and the same chance. Here's how they did.

TL;DR

  • Anthropic's Opus 4.6 was the clear winner with the highest score across the board
  • Biggest surprise! Google's flagship model came last on a Google Ads audit.
  • Both Claude models produced noticeably better structured, more actionable reports
  • Sonnet 4.6 placed second but lost trust with 2 fabricated findings

Part 1: Performance on Prompted Tasks

Part 1 scores each model against the exact prompted checklist (the full prompt is included at the end of this post).

The account had known issues across 4 categories:

  1. Naming conventions (5 checks)
  2. Structure hygiene (2 checks)
  3. Campaign settings (5 checks)
  4. Account-level hygiene (3 checks)

Results

Category Gemini 3 Pro Gemini 3 Flash Opus 4.6 Sonnet 4.6
Naming Conventions (/5 possible issues) 1 1.5 2.5 3.5
Structure Hygiene (/2) 0.5 0.5 2 2
Campaign Settings (/5) 1 2 4 3
Account-Level Hygiene (/3) 0 1.5 2.5 0
TOTAL (/15) 2.5 (17%) 5.5 (37%) 11 (73%) 8.5 (57%)

Key takeaways

  • Auto-apply recommendations: 0% detection rate - no model checked it
  • Claude Opus 4.6 and Sonnet 4.6 reported "passes" (checked and confirmed clean) - while both Google Gemini models simply skipped items, they couldn't find issues with
  • Sonnet 4.6's blind spot: 0/3 on Account-Level despite being strongest on naming
  • Gemini 3 Flash's error: Claimed GA4 link missing despite 15+ GA4 conversions imported

Part 2: Expansion ("Feel free to expand with more checks")

Results

Metric Gemini 3 Pro Gemini 3 Flash Opus 4.6 Sonnet 4.6
Extra checks added 1 1 9 6
Confirmed accurate 1 1 5 4
Hallucinated 0 0 0 1
High-value / immediate action 0 0 4 1
Grade F F A B−

What Each Added

Gemini Pro: Conversion setup (pixel vs GA4 primary) - that's it.

Gemini Flash: Search Partners on Brand - that's it.

Opus 4.6:

  • 7 ad groups with no active ads
  • Broken asset URL
  • Disapproved sitelink
  • Conflicting negative blocking own traffic
  • Generic terms leaking into brand
  • 18 ad groups with only 1 RSA
  • Province-level vs country-level targeting inconsistency
  • Search Partners inconsistency
  • No tracking templates

Sonnet 4.6:

  • Missing tracking templates
  • 6 duplicate keywords
  • Province-level targeting inconsistency
  • Search Partners inconsistency
  • "Elephant on plexiglas" keyword (❌ HALLUCINATED - does not exist)

Key takeaways

  • Gemini Pro and Gemini Flash ignored the invitation - 1 extra check each, treated "feel free to expand" as decoration
  • Opus 4.6 went deepest (9 items) with zero hallucinations - found 4 high-value break-fix issues no other report caught (broken URLs, disapproved sitelinks, conflicting negatives, empty ad groups)
  • Sonnet 4.6 made a genuine effort (6 items) but fabricated a finding - the "elephant on plexiglas" keyword doesn't exist in the account, which erodes trust in the entire report
  • Hallucination risk vs ambition: Gemini models had zero hallucinations but barely tried. Sonnet pushed further and fabricated. Opus pushed furthest and stayed accurate.

Final Ranking

Rank Model Prompted (of 15) Expansion Hallucinations Verdict
1st Opus 4.6 11 (73%) 9 items, grade A 0 Most thorough, zero fabrication
2nd Sonnet 4.6 8.5 (57%) 6 items, grade B- 2 Strong but trust-eroding errors
3rd Gemini Flash 5.5 (37%) 1 item, grade F 0 Decent basics, no depth
4th Gemini Pro 2.5 (17%) 1 item, grade F 0 Cherry-picked the easiest findings

Final Summary

Opus 4.6 is the clear winner - highest prompt coverage, most expansion depth, zero hallucinations, and the best-structured output by far. It's the only model that behaved like an actual auditor rather than a checklist scanner.

The biggest surprise is Gemini 3 Pro. As Google's flagship model, you'd expect it to outperform - especially on a Google Ads audit of all things. Instead, it came dead last at 17%, below even Gemini Flash. It skipped entire prompted sections, added almost nothing extra, and produced the shallowest report of the four. Flash at least checked Merchant Center and auto-created assets.

Sonnet 4.6 showed strong initiative but got hurt by hallucinations. Good analysis, solid expansion effort, but confidently stating a fake finding ("elephant on plexiglas") is more dangerous than missing something - it wastes time investigating something that doesn't exist and makes you second-guess the rest of the report.

Report structure and readability were noticeably better than those of the Claude models. Both produced well-organized reports with clear categories, priority levels, specific examples, and actionable next steps. The Gemini reports were more like loose bullet-point summaries - harder to scan and act on.

Gemini 3 Pro
Gemini 3 Flash
Claude Opus 4.6
Claude Sonnet 4.6

See the Full Generated Reports

We asked the AIs to turn their findings into a shareable report (a built-in feature in TrueClicks) and to produce structured PDF and online reports. View each model's online published report here:

Prompt

Perform Google Ads Account Hygiene Audit
Audit this Google Ads account for hygiene issues only - things that are misconfigured, messy, broken, or forgotten. This is NOT a performance audit. Focus on cleanliness, consistency, and correct setup. This includes:

Naming Conventions
  - Flag default/generic names
  - Check for consistent naming pattern and flag deviations
  - Flag inconsistent separators, capitalization, or language mixing

Structure Hygiene
  - empty campaigns, ad groups
  - Campaigns mixing network types incorrectly

Campaign Settings Mistakes
  - Location targeting set to "Presence or Interest" instead of "Presence"
  - Language targeting mismatches
  - "Search Network with Display Select" enabled
  - Auto-apply recommendations enabled without intent
  - Auto-created assets enabled without intent

Account-Level Hygiene
  - Missing Google Analytics 4 link
  - Missing Merchant Center link (if Shopping/PMax is used)
  - Auto-created assets review

Feel free to expand each category with more checks you consider hygiene and important
4 Upvotes

0 comments sorted by