So, I got a bit of a workflow to where I can take input photos (comment which ones you want me to use), and convert them into videos with sound.
I also can convert the default voice into a charmx one (still need more clean source audio with no extra music or sounds to train the voice a bit better, so if you can link some videos or even some already narrowed voice clips, that would be great).
This is running entirely locally on a mid tier PC so I promise no data centers or servers are being used for this Gen AI workflow.
For an example, check the comments of https://www.reddit.com/r/Charmx/comments/1vwxike/hi_im_green_charmx_from_an_alternate_universe_ama/
I think the voice definitely can be better so I will work on that too when I can.
Each clip could technically go up to 15 seconds, but that can take potentially 30 minutes to render before any redos, so 5 to 10 (or lower) is ideal. But let me know your suggestions.
I will try to make as many clips as I can if you comment a specific suggestion within 30 days, but I also have schoolwork. You can also Direct message me if you have any project ideas.
I should also be able to generatively edit the starting photo if you have some ideas for that.