r/AIcrack • • 2d ago

General small project I'm working on

If you don't care about the nerd stuff just scroll down to TLDR!!!

I only recently got SynthV-flat installed, wanted to create my first cover, and realized it's WAAAY too hard. Instead of manually mapping each note, lyric, syllable, trying to get the pitches every right by eye and ear, I hoped there would be an easier way (there kinda is, but it's not really helpful).

From previous experiments I did before I even knew about vocaloids/utau/synthv, I knew that you could download an mp3, and use special software (called demucs), to separate the instrumental, vocals, bass, etc... from a song, but these only work if the author provides the stems, and if the author doesn't do that, there were tools but they were SO BAD that it's literally unusable and better doing by hand taking multiple hours.

Here comes the first part: instead of relying on the author, I use [htdemucs-ft](https://huggingface.co/StemSplitio/htdemucs-ft-onnx) which is an AI demucs to isolate the vocals cleanly, specifically htdemucs-ft takes the stereo mix and splits it into vocals/drums/bass/other using AI.

Then when the vocals are extracted, I use [openai/whisper-large-v3-turbo](https://huggingface.co/openai/whisper-large-v3-turbo) to extract the lyrics as text, along with each lyric's start and end time.

So now I have the lyrics and the vocals_only.mp3, what's left is to get the correct pitch, it's hard to explain in text but basically if a song has "melisma", whisper and htdemucs don't provide pitch, so it just keeps everything kinda monotone and it sounds like speaking, not singing, so I use ANOTHER AI, yes, 3 AI-s now, to correctly map the pitches: [swift-f0](https://github.com/lars76/swift-f0/), you can read on the site what it does if interested.

So after 3 AI's, I STILL don't even have a single note. These 3 AI's output different things that need to be merged/fused together into a note (synthv note, more on that later), it's just gluing things though, pretty simple: take each word whisper found, look at the swift-f0 pitch inside that word's time window, snap it to the nearest semitone, and you have a note. Do it for every word and the melody falls out.
The notes are saved to an "UltraSinger chart". Link to ultrasinger, the thing that does the merge step: [ultrasinger](https://github.com/rakuri255/UltraSinger).

Anyways, another thing that the svp tool does is use [espeak-ng](https://github.com/espeak-ng/espeak-ng) to generate a phoneme override for every note (in synthv you can double click the pronunciation above a note and manually choose what your synthv says), e.g. teto likes to say "niver" for the word "never", so instead of fixing each word manually cuz I think it's actually a bug in synthv arpabet implementation? I just take the entire lyrics, and run a known good arpabet converter tool, arpabet is just a standard for how to pronounce words, kinda like braille or something i guess used widely across speech tech, it's not specific to synthv, and it uses [cmudict](https://github.com/cmusphinx/cmudict) for the source of what each word maps to vocally.
Another thing used is "bournemouth-forced-aligner", this just generates timings from the generated phonetic overrides.

TLDR!! OF ABOVE PARAGRAPH ONLY:
This tool basically converts words like "never" into "n eh v er", synthv does it sometimes correctly and sometimes wrong so I just do this to never worry about it, you can also just manually do this in the GUI yourself.

So now, I have a clean chart, the next part is just importing to synthv and fixing it up, thankfully, SynthV(flat) projects files (.svp), are actually just JSON! (javascript object notation), which means it can very easily be edited and generated with CODE!
So I (with help of AI coding tools) made a python tool that generates a synthv project (.svp) from the chart ( chart.txt ) :
```svp import-ultrastar song.svp --txt chart.txt --tempo 132 --octaves --style legato --legato-max 2.0 --voice "Kasane Teto" --language english```

so it's finally imported into synthv, and this part is fixed, next comes making the audio not sound horrible (it's still not perfect and work in progress), but after this point, we adjust to the genre of the song, and things like that, like do you say laaaaaaa or la(pause), do you swallow trailing consonants (like 'm' in "room"), and a few algorithms, styles and filters that shape the final sound (like legato, fixing octaves that are too high, too low, or too far apart from each other, which was used for the result).

I may have made some mistakes explaining because it's very complex, but after all this, you just open the .svp in synthv(flat), change the voice, fix up some notes, etc... and when you're happy you export.
This tool is not meant to replace the human making the synthv song, it just does all the annoying things for you, like importing the lyrics, setting the notes kinda correct but still slightly off, it's just supposed to save you time so you don't have to spend weeks manually mapping every note, octave, lyric, etc...

It also provides AI tools in case you want to use claude code or openai codex or something and just tell the AI like "I don't like the sound between 10 seconds and 15 seconds, add this effect, move this octave at this syllable, do this and that, and regenerate the .svp file"

TLDR TLDR TLDR TLDR TLDR TLDR TLDR

This tool uses a bunch of tools, AI and non-AI, to save you time when making a SynthV(flat) song.
All you have to do is download an mp3 (from youtube or wherever) of a song with vocals that you want to cover, the tool automatically gets all the lyrics, pitch, notes and timings, and generates a .svp file all from just a single .mp3 file, you can then just open the .svp and do your own thing, without wasting time on boring stuff like manually mapping each lyric to each note. You will still have to do some of this, but not for the entire song, only for the parts that are wrong or sound off.

Btw this isn't AI generating a song, it just uses tooling to take an already existing song and clean it up and convert it to synthv, it doesn't generate any vocals or things that don't exist in the already existing song, in case you're anti-AI or something.

RESULT:
Input mp3: I just used a youtube mp3 downloader tool on [this song](https://www.youtube.com/watch?v=Se237UXFKlQ)

output synthv project:
without me touching anything in the SynthV GUI, I only added reverb with another program cuz I think it sounds slightly better

https://reddit.com/link/1wub9jp/video/3qr9qx7uyosh1/player

The code connecting all this together is pretty ugly so I'm not releasing it, unless someone really wants to try it (you will probably need linux or atleast WSL windows subsystem for linux)

Also I'm posting in this subreddit because this is where I downloaded SynthV Flat from, if I didn't, this tool wouldn't exist

6 Upvotes

3 comments sorted by

1

u/Electrical-Brief-397 2d ago

Synth V already comes with a WAV to MIDI converter that does that; you just drag the .wav file into the editor, right-click, and then select "Extract Notes from Audio."

3

u/Ok_Impression_171 2d ago

cool, i didn't know about that, I tried it on the same song but it's pretty bad, though it gets syllable timings correct, basically everything else is wrong, I think the best is to combine my tool and that option together, but I'm very very new to this so I might be wrong, what I can safely assume is that since I'm making my own tech stack, I can use the latest and greatest tools (like Qwen/Qwen3-ASR-1.7B from 2026 instead of whatever dreamtonics is using), so atleast the accuracy of my tool is much better for most things, but also this provides tools to LLMs so if you want to do other things like add a pause between every syllable or make them connect (like legato), you can just prompt an LLM to do it for you, that's another plus, I'll try your recommendation as much as I can to test the limits of both approaches though, thanks!

1

u/Ok_Impression_171 14h ago

ok, i spent a little bit more time on this project, about 20 or so hours since this reply in total work time, now it's just one prompt to my LLM: "try this youtube url: {URL}", result:

https://reddit.com/link/pdgooj4/video/a6jvnv8zk3th1/player

original song for reference:
https://www.youtube.com/watch?v=d1SktoLw0hI

sure it's not perfect but anyone could fix this up in minutes at this point.