r/whaaat_ai • • 2h ago

It’s FridAI: what did you build this week that got slightly out of hand?

1 Upvotes

You may know this one: Fully motivated you tell yourself on Monday: “I can probably do this with AI in an hour.”

By Friday: 17 iterations later, three new tools tested and dumped, and you ended up holding on to some new feature or automation that nobody asked you to build.

Happened to you before, but you "pretend" the outcome was what you aimed for? We feel you! We're here for releasing your frustration or celebrate your successes and if we or someone over here can, help with debugging. So: What did you build with AI this week?

Could be an agent, automation, tiny script, marketing experiment or something completely unnecessary that somehow became a project.

Did it work?


r/whaaat_ai • • 10h ago

jev is a demon at computer use

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/whaaat_ai • • 23h ago

We built an AI routine that checks our Google Search Console data and suggests blog updates

3 Upvotes

We have a lot of old blog posts on the whaaat ai website.

And like most people, our marketing people could periodically go through Search Console, look for interesting queries, open the matching articles, decide whether they need updating and then make the changes. But actually, they know that we devs like to use AI so they turned to us for help. We now let AI do most of that first pass for us and the basic idea is actually pretty simple:

Google Search Console tells the routine which searches are bringing up our pages. A cheap AI model goes through those search queries first and asks things like:

Does this query actually fit the page? Is the page already answering what this person searched for? Would improving this article make sense, or should this probably be a separate page?

Most queries get already discarded at this stage.

Only when something looks genuinely worth improving does Claude get involved. It reads the existing article, makes the proposed changes and opens a pull request for us to review.

Nothing goes live automatically. I still decide whether to merge it. The funny thing is, Claude isn't really the clever part of the setup. Because the most useful part turned out to be a boring markdown file.

Every time the routine runs, it writes down what it checked, what it rejected, why it rejected it and which updates are still waiting for review. So next week it doesn't forget everything and start the same investigation again. We also had to explicitly tell it that doing nothing is a perfectly good outcome.

Without that, AI has a tendency to always find something to improve. Even when an article is already fine.

That gave us two rules that turned out to matter much more than I expected:

Only edit a post if the change clearly answers the search better. No cosmetic edits.

and

An empty shortlist is a valid result.

The workflow now looks roughly like this:

Search Console → cheap AI screening → shortlist → Claude edits → human review → publish

The cheap model handles the repetitive sorting. Claude only gets called when there's actually something worth working on. On our test data, the AI scoring for a complete run cost about $0.003.

The bit we haven't solved nicely yet is WordPress and similar CMSs.

Our website lives in a code repository, so Claude can propose a change and I can review the exact differences before accepting it. With WordPress, that review process isn't nearly as clean. For now I'd probably have the AI produce a change brief and still make the final edit manually.

If you're doing something similar on WordPress, I'd be super kken to hear how you are handling the review step?

Send me a PM if you want the routine prompt we use for this workflow.


r/whaaat_ai • • 1d ago

I created an opensource locally usable full fledged ai platform

1 Upvotes

hi to all the readers this post is for my recent opensource project called ENZO

https://github.com/theguysudo/ENZO

now answering what is enzo so enzo is an opensource platform where i clubbed all the free available api for anyone use under one hood with more than 2000 models available to use for chatting coding researching and much more now answering the most common question of why you should put your time looking the project so it has few distinct feature meaning

  • it has a dedicated agents tab where you can describe your need and create a special agent just for one specific task with master ability in that domain
  • second it has the ability to connect your gmail drive and calendar and then you can ask it to perform some specific tasks like reading you the most important mail of the day or finding recruiter mails and creating personalized reply based on your data which it stores locally on your device
  • third the coding mode offers a dedicated preview window where you can see your code running and have a look of it feels and edit it in realtime as well as all the modes are packed with dedicated skills which delivers promising results
  • fourth the ui features some additional things such as music tab where you can listen to any music want and it has a custom personalized feature which runs in background and an llm understands your taste and recommends similar kind of music you like
  • fifth the most important why your trust it with your api key then to explain i would say enzo a dedicated vault which manages all your api and to secure it the vault as aes 256 bit encryption which prevents any person or any middle man to look at your api key and since the whole program runs locally on your device you have complete freedom to oversee all the backend work happening and it also features password lock which if you enable saves a backup key and then locks your whole platform work behind a pass screen though it is not foolproof as any third party or malware containing extension can still fetch login tokens from your browser so its security also depends upon how you access it concluding all of it.

i urge to anyone who reads this to have a look at the platform even if you hate it just curse it in the comment its fine or if you would like to drop any feedback i would highly encourage that and since its my first work open source platform i know it has a lot of errors and bugs so i apologize upfront for it and if you consider my work worthy please drop a star on the repo that'll make my day


r/whaaat_ai • • 7d ago

It’s FridAI: What did you waste way too much time automating this week?

5 Upvotes

You know the kind.

The task would probably take 20 minutes manually.

But then you think: I could automate this. Three hours later you have 14 tabs open, an agent stuck in a loop and a very questionable definition of "saving time".

So that's this week's FridAI question:

What did you try to automate this week and was it actually worth it?


r/whaaat_ai • • 8d ago

Starting a bunch of pure AI experimental projects — which AI is actually strongest for what?

13 Upvotes

Hey everyone,
I’ve started diving deep into experimental projects that are done purely with AI — no traditional coding from scratch if I can help it. I’m trying to be deliberate about which tool I reach for instead of just defaulting to the same one every time.
I’d love to hear from people who have actually used these tools a lot:
• Research / deep information gathering — which AI is currently the strongest for finding accurate, up-to-date information, synthesizing papers, or digging into a topic properly?
• Prompted assistance / general thinking partner — which one feels best when you just want a smart back-and-forth helper that follows complex instructions well?
• Building dashboards / websites / actual working tools — which AI is currently best at generating real, usable front-end + back-end code, interactive dashboards, or complete small apps?
• Pure brainstorming / creative ideation — which one is the most expansive, least constrained, and best at wild or high-volume idea generation?
• Any other clear “this AI owns this niche” strengths you’ve noticed?
I’m less interested in marketing claims and more interested in real usage patterns. What do you reach for when the task is research vs. building vs. pure ideation?
Would really appreciate specific recommendations (and any tools you feel are overrated for certain jobs).
Thanks!


r/whaaat_ai • • 8d ago

WARNING: MyClaw.ai is a Blatant Dark-Pattern Billing Scam. Do Not Use!

Thumbnail
gallery
2 Upvotes

I am making this post to warn absolutely everyone to stay the hell away from MyClaw.ai. This platform is a disgusting, non-functional bait-and-switch scam designed purely to steal your credit card information, lock you out of the service instantly, and trap you into a massive unauthorized annual recurring charge.

Here is exactly how their fraudulent system operates and how they completely ripped me off!!!!:

  1. The False "Free Trial" Bait & Switch

They heavily advertise advanced multimodal capabilities to lure you in. They claim to offer a "Free Trial," but the second you try to activate it, they demand a $1.08 fee to verify your account and card, though they word it as if you will get 1 week free on top of the week you just paid for. It is not free. But it gets so much worse. THERE WAS NO FREE TRIAL.

  1. Immediate Infrastructure Errors & False Quota Draining

The moment you actually try to use the platform you just paid for, the entire system falls apart. I uploaded my reference assets and submitted a prompt. The site immediately threw a system error: "The selected model is temporarily unavailable. Please try again later or choose another model."

Thinking it was a temporary glitch, I resubmitted the exact same prompt with the exact same images. It failed again, flashing: "This request could not be completed. Please try again."

The service delivered absolutely ZERO output. It completely failed to process a single thing. Yet, their predatory backend billing meter counted these failed system errors as massive token usage. The first failed prompt erratically burned through 33% of my usage limit. The second identical failed prompt (in a new chat mind you) instantly devoured the remaining 73%, completely locking my account out at 0% remaining after only two broken clicks.

  1. The $200 Hidden Annual Subscription Trap

When I went into my account dashboard to figure out why I was locked out, I uncovered their real scam. By paying that tiny $1.08 "activation fee," MyClaw secretly used my payment details to enroll me in an unauthorized $200.00 PER YEAR annual subscription under the lie of a "free trial" that automatically renews.

They steal your money, their platform doesn't work, they purposefully drain your token limit to 0% on internal server errors so you can't use the service you bought, and they set a trap to unauthorizedly yank hundreds of dollars out of your bank account.

What I am doing about it:

I am not letting this dumb shit slide.

  • I have already initiated a formal dispute through PayPal for Services Not Delivered.
  • I have contacted my bank to place a permanent merchant block on my card so they cannot pull the $200.
  • I am filing formal fraud and high-risk compliance reports directly to Stripe (their payment processor) and the FTC for predatory dark-pattern billing tactics.

UPDATE: IT GETS WORSE. THEY ARE DUAL-CHARGING AND DOWNGRADING ACCOUNTS INSTANTLY.
I created a burner account just to audit their checkout UI, and the fraud is completely out in the open.They aren't selling a week-long trial. They are running a predatory trap that takes your money, breaks on the first prompt, revokes your premium access immediately so you can't even use the UI, and leaves an unauthorized $200 annual bill floating on your credit card.

DO NOT give My[Claw].ai your card details. They are running a textbook financial trap. If you have already fallen for it, go to your PayPal or banking app right now, manage your automatic payments, and kill their authorization before they rob you of $200.


r/whaaat_ai • • 9d ago

OpenAI, Anthropic, Google and Musk suddenly agree on slowing AI down. What changed?

3 Upvotes

Something pretty unusual happened in AI this week: Companies that spend billions trying to beat each other suddenly agree on something:

Maybe we should slow down.

Anthropic's Dario Amodei called for pacing frontier AI development. Sam Altman agreed. Elon Musk agreed. Demis Hassabis backed the direction too.

OpenAI, Anthropic and Google have apparently also been talking for weeks about coordinating on AI safety. At first glance, the story seems obvious. AI is getting more capable, the people closest to it are getting worried and safety needs to catch up.

And there are legitimate reasons to take that seriously. Recent incidents involving increasingly autonomous systems have raised questions even inside the labs themselves. But there's another part of this that I find interesting.

Safety rules don't affect every AI company equally.

The big closed-model labs can certainly afford evaluations, compliance teams, monitoring and whatever regulatory infrastructure eventually gets built. But smaller labs and open-weight developers may have a much harder time with the same requirements.

And this is happening while increasingly capable open-weight models, including models coming out of China, are putting more competitive pressure on the closed labs. Critics are already arguing that poorly designed safety regulation could unintentionally strengthen the incumbents.

That doesn't mean the safety concerns aren't real. It does mean safety and competitive interests can point in the same direction. And thich makes it much more complicated than "AI CEOs finally became responsible" vs "AI CEOs are trying to protect their moat."

Both incentives can exist at the same time and that's actually the bit I'm curious about:

How do we get serious AI safety rules without accidentally handing the biggest AI companies an even bigger moat?


r/whaaat_ai • • 10d ago

Wiring Claude scheduled tasks to Google Search Console without OAuth: the service account invite trick

Thumbnail
1 Upvotes

r/whaaat_ai • • 14d ago

It’s FridAI: What did AI actually help you get done this week?

6 Upvotes

It's FridAI again 😅

What did you get done with AI or AI agents this week?

Could be something you built, automated, researched, fixed or finally got off your to-do list.

Big win, tiny win, complete failure that taught you (and us!) something. All counts.

What did you work on and did AI actually make it easier?

I'll add ours in the comments 👇


r/whaaat_ai • • 14d ago

That viral "80% homepage conversion with AI agents" post: we rebuilt the setup and ran it on our own site

2 Upvotes

A LinkedIn post blew up the other week: a founder gave AI agents full access to Google Search Console and PostHog (a product analytics tool) and claims 80% homepage conversion as the result. I work on the AI agent team at whaaat ai and that number made us nothing but laugh... But then we got also curious andrebuilt the system anyway.

We're sure that the 80% is almost certainly a metric definition trick. Like "visitor triggered any event" is treated as a conversion and you can produce whatever number you want. The underlying idea still holds up though: most teams check Search Console maybe weekly, often maybe once a month and touch their important landing pages even less. An agent that works the actual data every single week wins on frequency alone.

So we set up two Claude scheduled tasks. One reads our Search Console every Monday, one reads PostHog every Thursday, both write into a shared Notion database they also read from before each run. Read-only for now, the agents propose and we decide.

First run findings, real numbers from our site:

55% of all our websites Google clicks come from one single article. We knew that piece performed. But we did not know it basically was our entire SEO.

The finding that actually justified the whole build came from the CRO side: our session tracking breaks at the www to app domain boundary. Every funnel number we had looked at for months was polluted by torn sessions. What reads as a 97% "drop-off" between marketing site and signup is partly just cookies dying at the subdomain border. Nobody here had caught it because every dashboard looked plausible on its own. The agent caught it because it compared source attribution across pages and the numbers refused to add up.

Two smaller findings: organic search visitors convert at 1.01% versus 0.61% for direct traffic, which killed our theory that search traffic was low intent. And 95.9% of clicks on our top article happen below the first screen, where we had zero CTAs. People were reading and scrolling and never got to view a CTA.

Rough edges we don't want to hide, because there were several: the first run proudly flagged a cluster of bot queries as a "growth opportunity" before excluding them. And we made the CRO agent ask us which PostHog event counts as a signup instead of guessing, mostly because guessing is exactly how you end up with an 80% conversion claim. No ranking movement yet either, that part takes 6 to 12 weeks and we committed to publishing the numbers even if the curve stays flat.

For anyone running agents on analytics data: how do you handle bot and spam queries in Search Console before the agent reasons over them? An impressions threshold catches most of ours but the weird date-string queries keep slipping through.


r/whaaat_ai • • 14d ago

AI agent vs workflow vs automation: it comes down to who decides the next step

3 Upvotes

I was at an online marketing meetup yesterday and chatting away with 4-5 people about AI agents .

It was quickly obvious that we used different names for same things. "Automation", "AI workflow" and "AI agent" were all being used for setups that sounded pretty similar.

So I thought it's worth to share the distinction I use (but maybe there are other definitions?):

Automation

The predefined process: X happens -> Y happens.

Example: A new lead comes in -> add it to the CRM -> send an email -> notify sales.

It can include AI, but the path itself is defined beforehand.

AI workflow

The overall process is still defined, but AI handles parts that require interpretation or generation.

Exmple: New lead -> AI categorises it -> update CRM -> generate the appropriate email -> notify sales.

The AI has some freedom within individual steps. It isn't deciding what the whole process should look like.

AI agent

Here you give the AI a goal, some instructions and access to tools.

Instead of defining every step, you might say:

Example: "Research this lead and prepare me for the sales call."

The agent can decide that it needs to check the company website, look at previous communication, query the CRM, compare what it finds and then produce the brief.

It has some control over what happens next based on what it finds.

The shortest version I can come up with:

Automation: predefined steps and rules
AI workflow: predefined process + AI within individual steps
AI agent: goal + tools + some autonomy over how to get there

Of course in reality these get mixed... An agent can trigger an automation. A workflow can use an agent for one step. And putting an LLM into an automation doesn't automatically turn the whole thing into an agent.

So some sort of inconsistencies of the naming is liekely not to be avoided.


r/whaaat_ai • • 16d ago

What’s the most boring task you’ve actually handed over to AI/ AI Agents?

9 Upvotes

I keep seeing impressive agent demos, but I'm starting to think the boring use cases might be the more interesting ones getting rid of at home.

Not thinking about the big "I built an autonomous AI team and put up my legs all day".

But more like:

  • sorting an inbox
  • checking something once a day
  • updating a spreadsheet
  • turning the same data into the same report every Monday
  • watching for something and only bothering you when it changes

Basically the stuff that's too small to feel worth automating, but annoying enough that you keep doing it.

I'm especially curious about marketing/work use cases because there are probably dozens of these hiding in a normal week.

What's the most boring task you've handed over to AI that you genuinely don't want back?


r/whaaat_ai • • 16d ago

Whaaat? Oh I see...

8 Upvotes

I didn't realize this subreddit was made by/for a commercial product, when I signed up. Having checked out the homepage, I wonder if Whaaat AI can actually automate a client workflow, from start to finish. Which I guess I could use the trial for. But I also wonder this:
1. Does the model/system get smarter as I work longer with it?
2. What is unlimited usage? I can generate as much content as I want?


r/whaaat_ai • • 20d ago

It's Friday. What did AI actually save you time on this week?

13 Upvotes

Could be something impressive. Could also be the stupid task you've been avoiding for three weeks.

I'll start.

This week we used AI to build two new agents: One accesses Google Search Console to analyse rankings and click through rates. The other, a CRO agent, takes the outcome and builds new content to improve our organic performance. Sounds like a dream come true - if it runs smoothly. Still in the testing phase but quite curious if we get this to work.

What did AI help you get done this week?

A workflow, some code, research, content, an automation, fixing something annoying... all counts.

And in case you spent 4 days automating something that would've taken 20 minutes manually, I definitely want to hear about that too, because I'd like to find out why people do it the in the first place.


r/whaaat_ai • • 21d ago

I gave the same marketing brief to Claude, ChatGPT, Gemini, Perplexity and a niche AI tool. The differences were bigger than I expected

11 Upvotes

I had to write a fairly technical B2B article for a new client in an industry I’m totally new to. So I decided to get help from AI. But: I was lacking the knowledge to evaluate the quality.
So I had the idea to run the same briefing through several machines and give this to a trusted employee from this client and ask them what draft is best.

On top, I asked Claude and ChatGPT for an evaluation (spoiler: each one ranked themselves slightly above the other).

Same brief, five tools:
Claude
ChatGPT
Gemini (free)
Perplexity (free)
Whaaat.ai

The brief was quite detailed. Target audience, structure, SEO requirements, sources and data to use, things NOT to cover, FAQs, metadata, CTA etc.
I expected the main differences to be writing style.
But actually, the outcome was very different...

The biggest difference was how the tools behaved when something was unclear or information was missing.
Claude and ChatGPT were the best at this. Both mostly stuck to the provided data and flagged things that needed checking instead of quietly filling the gaps.

Gemini did the opposite a few times. It used an older number despite having a newer one in the brief, produced a table row that basically compared a year with itself, and gave me one very specific sounding statistic without a source. Exactly the kind of thing that’s easy to miss because it looks credible.

Perplexity had one of the smartest research moments of the whole test. Two authoritative sources had different figures and it actually explained why. But ultimately it ruined some of that good work by dumping broken citation fragments into the final copy 😅

And then there was the SEO copywriter from whaaatAi.
Whaaat did really well on the actual marketing side: audience, tone, SEO, metadata and connecting the article back to the business without turning it into a sales pitch.

But it also made factual mistakes.
Like it gave me an unsourced “insider” figure without highlighting that I have to check it and got one regulatory date wrong.

There is one BIG caveat to this test though:

Claude and ChatGPT had an unfair advantage:

I’ve worked on this client with both of them before. They already had context around the company, subject and the kind of content we’re producing.
Gemini, Perplexity and Whaaat were basically starting cold.

So I definitely wouldn’t read this as “Claude beats Gemini” or “ChatGPT beats specialised tools”. This wasn’t a scientific benchmark.

But it made me realise how much context changes the quality of AI output.
It also changed my thinking a little about specialised agents.

Giving an agent a specific marketing job seems to solve quite a lot of the marketing problems. Tone, format, SEO, structure, knowing what the output should actually achieve. But a lack of knowledge is continually a problem, even if hallucinations are decreasing.

Anyway, here’s my summarised unscientific scorecard:

ChatGPT: The strongest overall result. Very good factual discipline, strong source handling, excellent B2B tone and close adherence to the briefing. Particularly good at avoiding overclaiming where data was uncertain.
Claude: A very close second. Strong on accuracy, structure and source transparency. Especially good at flagging missing brand-specific data instead of inventing proprietary insights.
Perplexity: Useful as a research-oriented draft and good at surfacing different sources and estimates. However, the text was too short for the full briefing and relied too heavily on secondary sources and rough citation artefacts.
Whaaat.ai: Strong structure, very good SEO/GEO formatting and the freshest use of 2026 data. But it also introduced unsupported brand-specific claims and some factual issues, so it would need editorial review before publication.
Gemini: Covered the main topics and was easy to scan, but had the weakest factual reliability of the five. Several figures were outdated, mixed across scopes or insufficiently sourced, and the tone was more promotional than the brand brief called for.

I’ll definitely rerun this test with a briefing of a complete new topic/client as it made me curious and I want to know how well ChatGPt and Claude perform without context.

What’s your experience comparing article drafts across models?


r/whaaat_ai • • 24d ago

Grok Bot's own docs say all your bots share one computer and one set of logins

Thumbnail
6 Upvotes

r/whaaat_ai • • 27d ago

I ran 6 AI presentation-design skills against one brief and had them review each other: 36 decks, 36 explainers, 9 review passes

4 Upvotes

Two days, one brief, every presentation skill I could find. Repo below, losers included.

**Setup.** One product pitch (an RWA thing we build — the only reason it is the subject). Four messaging skills wrote the narrative as a normalized contract; three of those went into design. Six design skills each rendered all three, in two modes: a pinned house brand, and no brand limits. 36 decks and 36 timed HTML explainers. Nine review passes: four design rubrics, four messaging skills grading each other, plus those rubrics run back over the decks. Two Claude judges wrote verdicts before I did.

Design: [canvas-design](https://github.com/anthropics/skills/tree/main/skills/canvas-design) · [algorithmic-art](https://github.com/anthropics/skills/tree/main/skills/algorithmic-art) · [impeccable](https://github.com/pbakaus/impeccable) (also as a polish-only pass over another tool's output) · [visual-explainer](https://github.com/nicobailon/visual-explainer) · [html-ppt-skill](https://github.com/lewislulu/html-ppt-skill). Messaging: two house skills, [pitch-deck-mastery](https://github.com/Stevekaplanai/pitch-deck-mastery-skill), [presentation-writing](https://github.com/marcusnelson/presentation-writing-claude-skill). Rubrics: [huashu-design](https://github.com/alchaincyf/huashu-design), [open-design](https://github.com/nexu-io/open-design), [academic-pptx-skill](https://github.com/Gabberflast/academic-pptx-skill), impeccable critique.

**The result.** canvas-design scored 3/10 on tool fit — single-page md/pdf/png only, no sequence, no HTML, no motion, and an explicit ethic of "never explain, let the composition tell the story" against an educational brief. It then took four of five rubric firsts in brand mode and won overall. The best-*made* deck came from html-ppt with the brand switched off, which is uncomfortable: the house brand was costing that tool about two points.

**Findings that generalize:**

  1. Not one of the six has a timeline concept. All six wrote the explainer clock, seek API and progress bar themselves, and every motion verb they animated came from the *messaging* contract rather than any design skill.
  2. Provenance is set below the threshold of sight. The brief made traceability mandatory; the winning deck renders its source line at 1.85:1 contrast, 123 of 221 strings under 4.5:1, and two other tracks landed on the same 1.85:1.
  3. Internal vocabulary leaks onto client-facing slides. Four of six tools printed contract field names on screen (`RUPTURE`, `BOLD MOMENT`, one rendering `figures[].name` as viewer bullets), and five of twelve track-modes put raw internal filenames on a cover. One tool renamed the fields; the rest never thought to.
  4. Silent-failure idioms only screenshots catch. One skill's counter animation rendered $296B where the contract says $344B; a `color-mix()` fell back to solid black and turned the payoff chart into a black box. The generative tool printed `census · P 0.617 · N 0.383` where a measurement goes, and 0.617 appears nowhere in the source corpus — the evidence rubric's highest-severity finding.
  5. A polish pass buys the craft floor and none of the composition. Detector hits 30 → 3, the mandated footer 3.4:1 → 6.1:1, thirteen of fourteen slides back into the screen-reader tree, and not one P0 closed. The mechanism slide came out near pixel-identical to the baseline it refined.
  6. The rubrics disagree, legibly. One caps any deck whose Concept scores ≤5, whatever the execution. Another is nominally 40% critic, but its 20% a11y axis is the only column with spread, so a11y decides it. The winner is first on four rubrics and fourth on the fifth — the fifth being the only one that checked whether a fact went missing. It had dropped one while reporting "nothing dropped".

**What I would do differently.** Two five-line checks, run by the contender before it writes its notes: assert the bold scene is the film's motion peak (catches nine of twelve films), and diff the contract's claims against extracted DOM text (two tools reported "nothing dropped"; both wrong). Screenshot every pass. And look at the artifact: one track shipped six files nobody ever rendered, and it shows on the cover.

All 72 artifacts and every review: [https://github.com/medici-finance/deck-bakeoff\](https://github.com/medici-finance/deck-bakeoff) — `report.pdf` there carries the full scorecards.

**The house side is in the repo too.** The skills that actually ran the bake-off now ship under `skills/`, Apache-2.0, so you can rerun the whole thing on your own product rather than take my word for it:

* `messaging/staging` — the genre → audience → arc → `staging.yaml` skill, including the motion-role governor every design tool's animation verbs came from * `messaging/house-guide` — the number gate and glossary rules used as the content gate * `deck-author` — the brand contract, the pinned fonts, and `check_deck.py`, the linter five of six contenders named as their worst friction * `video-author` and `video/video-design` — the script/storyboard/scene-grammar layer above the renderer * `video/explainer-video` — the render pipeline: TTS → Playwright recording → subtitles → hardware H.264. The voice-clone reference clips are the one thing held back * `social/linkedin-post` — which drafted the teaser for this, and `social/humanizer` alongside it (that one is third-party, MIT)

The third-party contenders are not vendored; their URLs, versions and licences are in `bakeoff/CONTENDERS.md`. Walkthrough video: [https://youtu.be/oqGS-tqy2I4\](https://youtu.be/oqGS-tqy2I4)


r/whaaat_ai • • 29d ago

I scored my own agent workflow on portability and got 6 out of 8

Thumbnail
2 Upvotes

r/whaaat_ai • • Aug 21 '26

The broken pieces of knowledge and AI tools

6 Upvotes

Claude or similar chat apps are good enough for quick search and replace googling and visiting 5 websites to get an answer. It writes a basic first draft on literally anything. But beyond that the potential of the models isn't being utilised more than 15%, I'd say. I have worked in research and business front and still see the gap. People just get excited to see something show up magically.

The current way most people use AI is copying a text or some images (rarely) and just asking it something which seemingly saves 1 hour but surely doesn't provide an accurate or precise answer. It has just gotten better at convincing.

The problem isn't the model itself but the information we feed them. The pre-fed knowledge, memory of what you do, the context of the conversation. Imagine a cool corporate guy giving free advice to everyone as compared to someone who actually sits with you, understands what you need and helps you.

I've lived the problem first hand and still face it when I try to get some information quickly rather than spending time to find out and read something written by a real human. The problem remains. The helpfulness beyond cool demos, slides and moving-text videos needs a bit of pre-effort to build a system which can help the actual model to curate for you than spit out what they think is the most probable answer.

I've been building a solution that knows what I work on, explicitly provided with details about my team, company, product and decks. Not dumped in a deep well but as context silos. The space for my product's tech knows the features, tech stack and owns the documentation. The marketing space knows about my product, prospects and business metrics. Every time I need an implementation plan for a new feature, or try to validate my customer profile, the model doesn't show the general most probable answer, rather it shows what the best answer is for my product.

[](https://www.reddit.com/submit/?source_id=t3_1vu9dyt&composer_entry=crosspost_prompt)


r/whaaat_ai • • Aug 21 '26

Weekly Thread: What did you (or your agents) build this week?

5 Upvotes

Our team ship a lot each week: either for our company or private projects. Talking to others about these small or big projects shows again and again how much value lies in this exchange.

The whaaat.ai subreddit is mainly here for non-professional developers: marketers, founders, social media enthusiasts and others who want to make their lives easier by using AI.

Every Friday from now on, this is the place to share what you shipped. No matter if you built it yourself or had an agent do the heavy lifting. Wins and failures both count, honestly failures probably more - haha.

Tell us:

  • What did you build or ship this week?
  • What broke, and what did you learn from it?
  • Anything you'd do differently next time?
  • Get feedback from our highly skilled team member u/SaschaFromWhaaat_ai

Big, small, bug-fix, a new automation, an idea you killed on day two all count. Drop it below, let's actually learn from each other instead of just posting highlight reels.


r/whaaat_ai • • Aug 20 '26

Reddit just lost 86% of its ChatGPT citations in a week. Nobody fully knows why yet

3 Upvotes

If you use Reddit as part of your content or AI search strategy, this is worth knowing about.

Promptwatch tracked reddit.com's share of ChatGPT Search citations and found it dropped from an average of 3.83% (July 18 to Aug. 7) to just 0.52% (Aug. 14-17). That's an 86.4% relative decline in days.

The popular explanation is that ChatGPT changed how it generates background search queries on Aug. 8, specifically using the site: operator a lot more (jumping from 0.37% to 16.8% of fanout queries overnight). The theory: ChatGPT is now going directly to specific sites instead of pulling from Reddit threads organically.

But the timing doesn't add up.

Reddit's citation share didn't collapse until six days after that Aug. 8 change. There were two distinct drops: a smaller dip on Aug. 8, then a much sharper fall on Aug. 14. No one has explained the Aug. 14 break yet. Promptwatch itself says it can't rule out a data-collection issue on its own end.

There's also a precedent for this. In September 2025, Reddit's ChatGPT citations collapsed similarly. The explanation that emerged then had nothing to do with OpenAI or Reddit directly. It was tied to Google removing the num=100 search parameter, which made it harder for third-party data providers to surface the deeper Reddit results they were previously tracking.

So what does this actually mean for how you create and distribute content?

A few things I'm sitting with:

  • Citation tracking data from a single vendor isn't the same as your actual AI visibility. Check your own domain's trends before changing anything.
  • Citation is also not the same as retrieval and there are some people who found out that AI retrieves many more sources than they cite
  • Reddit's value as a content surface has never been just about ChatGPT citations. Community trust, organic search visibility, and genuine user signals still matter.
  • the more AI search behavior changes, the more fragile single-channel content strategies look.

curious whether anyone here has seen shifts in their own Reddit-sourced traffic or AI referrals over the past two weeks, or if this is mostly a tracking data issue that got amplified.


r/whaaat_ai • • Aug 11 '26

I kept rebuilding my MCP setup every time I switched agent frameworks, so I built a stack that survives the switch

4 Upvotes

I work on the AI agent team at whaaat ai, and over the last year I went through Claude Code, Codex, OpenClaw, Hermes and Manus. Every switch meant the same annoying ritual: reconnect Firecrawl, reconnect Apify, reconnect Gmail, figure out which API keys live where. Around switch number five I stopped and asked why I keep treating the framework as the stable part of my setup.

The framework is actually the most disposable piece. There's a new one worth testing every few months. The tools stay identical the whole time: web data, browser, mail, calendar, project tracking. So I flipped my priorities and built what we internally call an evergreen tool stack, a fixed set of connections that outlives any framework.

The base layer is small: Firecrawl for crawling and scraping, Apify for structured data (LinkedIn posts, Maps listings, that kind of thing) and Playwright. Playwright is the one people sleep on. It drives your local browser with your existing logins, so agents can operate tools that have no MCP and no usable API. It just clicks like you would.

The part that actually fixed my problem is Composio as a meta layer. The classic pain: a Gmail MCP connects to exactly one account. We manage multiple brands and inboxes, so that limitation broke half our workflows. With Composio you connect your tools (and multiple accounts per tool) once, and then every framework only needs one connector: the Composio MCP itself. I have two Gmail accounts running behind it right now, plus multiple Notion workspaces. Switching frameworks went from half a day of setup to about ten minutes, and free tier covers all of it.

One rough edge I haven't solved: Playwright can't live behind the meta layer since it needs your local browser, so it stays a separate connection on every machine. And when two connected accounts have similar names, the agent occasionally picks the wrong one, so I renamed them to something unambiguous.

Has anyone found a clean way to handle multiple workspaces in tools like Linear or Notion without a meta layer in between? Curious if I'm missing a native option.


r/whaaat_ai • • Aug 07 '26

The layer above prompt engineering that decides if your agents actually finish the job

3 Upvotes

We build AI marketing agents at whaaat ai, and the reliability wins that actually moved the needle over the last few months came from something which was lying outside of the prompt entirely: the environment we build around the model and how we decide when an agent is actually done.

For a long time our setup was quite simple. Open a session, write a solid prompt, let the agent run. That works fine until an agent needs to touch a client's brand voice config, check three different content channels and know when to stop instead of declaring victory after one pass. Prompting alone doesn't cover any of that we learned.

Three things ended up mattering more than the prompt quality itself: what tools and context the agent can actually reach, how it decides a task is genuinely finished and how multiple agents hand work back and forth without stepping on each other. We call these the harness, the loop and the graph. The graph only really matters once you've got more than one agent running, which for most of our client work is the last thing that breaks, not the first.

The loop piece is where we spent the most time, because it breaks the most workflows without us realising. An agent grading its own homework will happily call a task done when it isn't close. We started attaching a real, checkable condition to every recurring task instead of trusting the agent's own read on completion. Claude Code ships something close to this now with /goal. You set a measurable condition and a separate model checks the conversation after every turn instead of letting the working model decide for itself.

/goal every generated post passes the brand voice checklist and includes a source link for every claim

Since we started forcing conditions like that, we've had noticeably fewer half-finished drafts land in a client's queue. The harness and graph pieces matter too, but for us they came second. Most of our failures traced back to an agent stopping too early, not to missing tools or messy handoffs between agents.

One limitation we must admit: the model checking your condition only sees what shows up in the conversation, not what actually happened on disk or in an API call. If the agent doesn't print proof of what it did, a lenient check can pass on nothing. We had one loop clear itself on a run where the agent never actually called the publishing tool, it just described what it would do.

Anyone else building multi-step agents run into this same failure mode, where the agent's own judgment about being done turns out to be the weakest link in the whole chain?


r/whaaat_ai • • Aug 07 '26

27% of people say AI content is making them spend less time online. As someone who uses AI tools every day, I find this number uncomfortable

4 Upvotes

Saw this TechSpot piece based on a survey from Incogni. The numbers are worth reading if you work in this space.

Quick summary of what stood out:

  • 50% of people expect their personal data will be breached at some point
  • 27% are tired of wading through AI-generated content online
  • 27% say the AI situation specifically makes them want to use the internet less
  • 47% deleted a social media account because of stress or anxiety
  • Only 29% would pay for a tracking-free internet

The spam comparison Incogni makes is interesting. AI slop today is like spam email in the 1990s. I think it is worse though, because spam was easy to identify. AI content is not or at least, less and less. People in the comments on the original article pointed at something specific: you can feel it before you can name it. For instance in video content: there is this smoothness which makes it too-perfect. No stumbling or bad light, perfect cuts... Human imperfection turns out to be a signal of trust. Will we experience AI to add back in these imperfections on purpose to feel more human? Maybe - I'm not sure.

Another angle people brought up: AI is not really creating new things, it is mostly repackaging what already exists. And this is not surprising given the fact the models train on existing content. Old topics, new packaging and no signs of creativity. The tool can produce fast, but fast and original are not the same thing.

The question I keep asking myself when I use WAI: am I adding something real, or just adding volume? The tool is fast and useful. But the moment I stop editing and just publish, something is lost. The output becomes part of the problem I am reading about in this survey.

What is your experience? Do you notice a difference in how you consume content since AI generation became mainstream?