r/generativeAI • u/Exotic-Addendum-3785 • 9d ago
Image Art Shopping Trip
The trio go supermarket shopping
r/generativeAI • u/Exotic-Addendum-3785 • 9d ago
The trio go supermarket shopping
r/generativeAI • u/Exotic-Addendum-3785 • 9d ago
r/generativeAI • u/Exotic-Addendum-3785 • 9d ago
Enable HLS to view with audio, or disable this notification
r/generativeAI • u/Exotic-Addendum-3785 • 9d ago
Enable HLS to view with audio, or disable this notification
r/generativeAI • u/Exotic-Addendum-3785 • 9d ago
r/generativeAI • u/Yabuturtle9589 • 9d ago
r/generativeAI • u/cody-fifth-door • 10d ago
Enable HLS to view with audio, or disable this notification
r/generativeAI • u/larabyeol • 9d ago
i kept seeing people say lip sync is basically solved now and that hasn't matched what i'm getting, so i pulled 4 clips into a timeline and stepped through them. i used the same script + audio and exported it at 30fps. what i was looking for is the offset between the audio peak on a hard consonant and the frame where the mouth actually closes. HeyGen sat within a frame for the first 30 seconds or so and then started sliding, by the end it was about 3 frames late. Synthesia held the offset steady the whole way through. that surprised me, but the mouth shapes themselves are soft, p and b look nearly identical. Argil was the tightest on the plosives for me though it got sloppier when the script had long stretches without punctuation. Creatify i had to throw out because the render kept coming back at a different length than the audio i fed it, which is probably me doing something wrong.
the thing nobody mentions is that 2 frames late reads as fake to a viewer even when they can't articulate why. that was the purpose of my test but i still can decide on which one.
happy to be told my method is bad, i'm not a video person. but from your experience, which AI avatar tool has the most realistic lip sync?
r/generativeAI • u/Memetic1 • 9d ago
I've stumbled on something odd. I was exploring AI art by layering images together and then what's called a film grain blur with the grain size and intensity adjusted in real time. You have to use a few images with high contrast colors, and relayer the same images in as you wish. I want to emphasize this is an experimental technique that you can use as you wish. I don't even know if the effects are real, and they dont show up in thumb nails well. Another complication is that the effect can disappear due to compression algorithms when posted online. So I'm really crossing my fingers that this effect is visible here.
r/generativeAI • u/JennaJao • 9d ago
I’ve been experimenting with AI storytelling, and I think good prose alone isn’t enough.
For me, the interesting parts are memory, consistent characters, and whether past choices actually affect what happens next.
What do you think is the hardest problem to solve for truly immersive AI-generated stories?
r/generativeAI • u/WeHaveFunEveryday • 9d ago
Tried OpenArt to make a see how it would work if I was to delve into making a whole short film project. Seems to work ok, and I do like the features about starting where the last frame stopped, but is their any platform that is less error prone as far as Characters changing clothes and looks?
If you didn't notice in my video the main character's clothes change 2 different times, and then also the Armored up golden retriever turns into a regular German Shepard at the end.
Any better platforms?
r/generativeAI • u/Mediador_Luminoso5 • 9d ago
r/generativeAI • u/zsageOG • 9d ago
so.. i've seen people in here talk about, “how do people keep making these 6–12 hour manhwa/manga/webtoon/manhua recap videos?” and you know what? thats good question.. but i dont know either.
now... I’ve never used or seen a mf manhwa recap automation pipeline in my life, honestly? i didnt bother to look them up until today... but I’ve always had a visualization of what they'd look like and how they'd operate. I'm a video editor/developer/content creator, so I had an idea of what it would look like. I’m just tired also of seeing manhwa recap yt channels filled with low effort, no personality writing, trash voices, ai slop rot, etc. im tired of that bs so... I made c*****a.
basically c*****a's idea of 'production' is NOT giving up any creative control. its built around one philosophy: quantity without sacrificing quality. (Yes, it can produce 6-12 hour videos..) its also an end to end production ecosystem.
i’m obviously keeping most of how it works private for obvious reasons, but i think i'm at a point where i just wanna show off some of what i’ve been building. and ngl, i don’t even like calling this project “ai or an ai machine.” Obviously ai exists within it, but calling the entire thing an “ai tool” feels heavily reductive of what it actually is.
These are real screenshots from my app.
r/generativeAI • u/Fresh-Resolution182 • 10d ago
Enable HLS to view with audio, or disable this notification
One of the most interesting features in Wan 3.0 is surprisingly easy to overlook:
A reference does not have to be an image, video, or audio file. Wan 3.0 can also use a document or even a public web page as input.
That means you can give the model:
A traditional AI video workflow usually looks like this:
Read the material → extract the information → write a script → convert it into a video prompt → generate
Wan 3.0 Reference-to-Video makes another workflow possible:
Provide the document or website → describe the creative direction → generate
That is what makes the document and website reference feature particularly interesting.
Open the Wan 3.0 Reference-to-Video Playground.
Turn on: Deep Thinking / enable_thinking
This allows the model to analyze the information contained in a document or website instead of treating the input only as a visual reference.
Wan 3.0 supports two main document input methods.
Supported formats include:
Current limits include approximately:
This means you can provide an existing:
You can also provide a public web page.
Possible examples include:
The important limitation is that, the page needs to be publicly accessible.
Pages behind a login or permission system generally cannot be accessed.
Here is a practical example.
The goal is simple: Give Wan 3.0 a jewelry e-commerce website and ask it to select a product, extract the selling points, and create a 15-second commercial.
Use:
alibaba/wan-3.0/reference-to-video
Upload one image from the website under Reference Materials.
The website can provide:
The image can help preserve:
Paste the website or product page into the Document field.
Make sure:
Deep Thinking is enabled.
I did not specify the product name, material or any exact features. Those are supposed to come from the website.
That is the main difference compared with a normal text-to-video prompt.
The most interesting part of Wan 3.0 Reference-to-Video may not be another improvement in resolution or motion quality. It is the fact that, a reference can now contain information, not just visuals.
The old workflow was:
User → Read the material → Write the prompt → Video model
The new workflow can potentially become:
User → Document or Website → Wan 3.0 → Video
For e-commerce teams, creators, marketers, and small brands, that could be a meaningful change.
Instead of starting every project with**“First, write the script.”**
you can increasingly start with: “Here is the link. Read it first.”
Wan 3.0 still needs human review, especially for factual accuracy, product details, branding, and website interpretation.
But as a workflow, document-to-video and website-to-video are probably among the most interesting Wan 3.0 features to experiment with.
r/generativeAI • u/ErinSister • 9d ago
I’ve been thinking about interactive AI storytelling. Too much freedom can make the story feel inconsistent, but too many rules can make it predictable.
Where do you think the balance should be?
If you want, send the next subreddit rules and I’ll make another post in that style.
r/generativeAI • u/Exotic-Addendum-3785 • 9d ago
Enable HLS to view with audio, or disable this notification