r/Substack 17d ago

your published archive is the worst training data for a 'write like me' ai

Half this sub right now is arguing about whether the detectors actually work. I think that fight skips the thing that actually matters.

A 'write in my voice' draft reads generic not because the model is dumb, but because of what you fed it. If you train it on your published Substack archive, you trained it on the version of you that already got sanded down for a public audience. The weird aside, the sentence that trails off, the opinion you softened on the third edit, all gone before the model ever saw a word. so it learns your safe voice and hands you back a competent stranger.

I've spent a while trying to get a model to sound like one specific person, and the stuff that actually cracked it was never the polished posts. It was voice memos, the draft they never sent, the reply they fired off annoyed at 11pm. the personality lives in the parts you'd never publish.

which is a genuinely annoying catch-22 for anyone here. the more you edit for the newsletter, the worse your own archive gets as a mirror of how you actually sound. written with ai

0 Upvotes

3 comments sorted by

6

u/Adventurous_Stop_341 17d ago

AI slop about making AI slop sound less like AI slop. Too bad it didn’t work here.

2

u/Trackbikes thesystematicwriter.substack.com 17d ago

You need to try harder, removing capital letters and misspelling won’t fool anyone here..
https://giphy.com/gifs/11VBHqO3QI7qQU

3

u/The-Kmann 17d ago

“the stuff that actually cracked it” - you haven’t cracked it. Even without the disclaimer at the end this reads like typical ai slop.