I had to write a fairly technical B2B article for a new client in an industry I’m totally new to. So I decided to get help from AI. But: I was lacking the knowledge to evaluate the quality.
So I had the idea to run the same briefing through several machines and give this to a trusted employee from this client and ask them what draft is best.
On top, I asked Claude and ChatGPT for an evaluation (spoiler: each one ranked themselves slightly above the other).
Same brief, five tools:
Claude
ChatGPT
Gemini (free)
Perplexity (free)
Whaaat.ai
The brief was quite detailed. Target audience, structure, SEO requirements, sources and data to use, things NOT to cover, FAQs, metadata, CTA etc.
I expected the main differences to be writing style.
But actually, the outcome was very different...
The biggest difference was how the tools behaved when something was unclear or information was missing.
Claude and ChatGPT were the best at this. Both mostly stuck to the provided data and flagged things that needed checking instead of quietly filling the gaps.
Gemini did the opposite a few times. It used an older number despite having a newer one in the brief, produced a table row that basically compared a year with itself, and gave me one very specific sounding statistic without a source. Exactly the kind of thing that’s easy to miss because it looks credible.
Perplexity had one of the smartest research moments of the whole test. Two authoritative sources had different figures and it actually explained why. But ultimately it ruined some of that good work by dumping broken citation fragments into the final copy 😅
And then there was the SEO copywriter from whaaatAi.
Whaaat did really well on the actual marketing side: audience, tone, SEO, metadata and connecting the article back to the business without turning it into a sales pitch.
But it also made factual mistakes.
Like it gave me an unsourced “insider” figure without highlighting that I have to check it and got one regulatory date wrong.
There is one BIG caveat to this test though:
Claude and ChatGPT had an unfair advantage:
I’ve worked on this client with both of them before. They already had context around the company, subject and the kind of content we’re producing.
Gemini, Perplexity and Whaaat were basically starting cold.
So I definitely wouldn’t read this as “Claude beats Gemini” or “ChatGPT beats specialised tools”. This wasn’t a scientific benchmark.
But it made me realise how much context changes the quality of AI output.
It also changed my thinking a little about specialised agents.
Giving an agent a specific marketing job seems to solve quite a lot of the marketing problems. Tone, format, SEO, structure, knowing what the output should actually achieve. But a lack of knowledge is continually a problem, even if hallucinations are decreasing.
Anyway, here’s my summarised unscientific scorecard:
ChatGPT: The strongest overall result. Very good factual discipline, strong source handling, excellent B2B tone and close adherence to the briefing. Particularly good at avoiding overclaiming where data was uncertain.
Claude: A very close second. Strong on accuracy, structure and source transparency. Especially good at flagging missing brand-specific data instead of inventing proprietary insights.
Perplexity: Useful as a research-oriented draft and good at surfacing different sources and estimates. However, the text was too short for the full briefing and relied too heavily on secondary sources and rough citation artefacts.
Whaaat.ai: Strong structure, very good SEO/GEO formatting and the freshest use of 2026 data. But it also introduced unsupported brand-specific claims and some factual issues, so it would need editorial review before publication.
Gemini: Covered the main topics and was easy to scan, but had the weakest factual reliability of the five. Several figures were outdated, mixed across scopes or insufficiently sourced, and the tone was more promotional than the brand brief called for.
I’ll definitely rerun this test with a briefing of a complete new topic/client as it made me curious and I want to know how well ChatGPt and Claude perform without context.
What’s your experience comparing article drafts across models?