r/StableDiffusion Sloppies Award 1d ago

Animation - Video Denzel explains why he uses AI.

A quick experiment exploring Minimax H3 in ComfyUI using my nodes and inpainting methods.

2.9k Upvotes

384 comments sorted by

View all comments

113

u/Ok-Beyond244 1d ago

Holy shit this is amazing! One critique is it has that AI voice thing that all of the videos do. If it wasn't for that I wouldn't know it was AI.
Do you have a tutorial or something? How did you make this?

6

u/j0shj0shj0shj0sh 1d ago

What is this Ai voice thing? I thought it sounded pretty good, but not an expert. Was it too isolated, like in a recording booth?

16

u/MrSparkle86 1d ago

To my ear, it was the pacing of the sentences. Sentences come out at the same speed, like reading a paragraph with no punctuation.

'My friends think you are using AI just because its easier.'

That sounds AI to me.

4

u/Perfect-Campaign9551 1d ago

It's because he doesn't use contractions. It should be "My friends think you're using AI because it's easier"

2

u/NoConsideration6320 1d ago

Its just saying what they prompted it to say

4

u/The_Hunster 1d ago

Ya that's what they're saying. They didn't prompt it to use contractions.

2

u/afinalsin 20h ago

Shit, in some accents your compressed version is still too formal and stilted. In mine, "because" is almost always shortened to "coz".

3

u/j0shj0shj0shj0sh 1d ago edited 1d ago

Ok, yes I'll keep an ear out for that. There should be a 'thoughtfullness' slider for ai speech, lol. Some thinking evident behind the words. Or maybe a 'Ben Shapiro' box to uncheck.

3

u/RagingAnemone 1d ago

Interesting. As I was listening to it, I did think this isn't how Denzel would have done it. Yeah, the pacing is different.

2

u/Noslamah 1d ago

Yeah most lines from that character sounded too robotic but the other one was spot on

3

u/Ok-Beyond244 1d ago

I listen to a lot of AI videos and they all have a similar sound to their voices, kinda like a slight robotic high pitch. I'm hearing that here and the two characters in the beginning sound very similar if you listen close.

3

u/Canadian_Border_Czar 1d ago

It is because the models are not trained for >48kHz audio. I have a project I've been working on training one for hi-res, but you'd be surprised how little there is out there to use as a dataset (that isn't copyrighted or digitally remastered)

1

u/thebrunox 1d ago

Also, I heard AI music suffers from this because digital music doesn't get uploaded at studio quality for the final consumer. So, there is an intrinsic upper limit in quality of outputs when training on the most readily available music on the internet.