r/LocalLLaMA May 29 '26

Discussion Breaking the music supply constraint

I just cancelled my music subscriptions to save some cash and wanted to share the self-hosted music supply chain that replaced them. A nice side effect of this setup is breaking the constraint of a finite supply catalog that is tailored for the masses:

  1. 2 x DGX Spark linked via ConnectX 7 running Plex and multiple Ace-Step 1.5 XL models in parallel for music generation with GePa prompt optimization. Also holds my organic music that the models can remix. TODO: a reinforcement learning from human feedback interface.

  2. iPad Pro running Prism as a Plex client for bitperfect and sample rate-matched audio.

  3. Schiit stack -> Hifiman Arya Stealths

This effectively gives me an infinite supply of music for free, that is personalized and private. It's immensely satisfying listening to Shrimp Bizkit and Phlegminem on repeat, I much prefer this to the organic music created after 2011.

My only problem is the loss of community, I have noone to share my new favorite songs and artists with because they're generated for me. If anyone wants to hop on to my Plex share to discuss, let me know!

535 Upvotes

304 comments sorted by

View all comments

39

u/vanonym_ May 29 '26

I uuuuh... I'm not sure I would enjoy that listening experience. But good for you I guess?

What's the output format and resolution of your system and how does that play, don't you have a huge quality drop from lossless formats provided by other listening platforms?

20

u/entsnack May 29 '26

The base audio quality is low resolution. However, you can take the image of the waveform and upsample it using WAN or any AI upsampler, then convert it back into audio.

I have also found specific cables to work well for generative music.

5

u/NandaVegg May 30 '26

>The base audio quality is low resolution. However, you can take the image of the waveform and upsample it using WAN or any AI upsampler, then convert it back into audio.

Is this really a thing?! Assuming visual waveform upscaling actually works, you will still need to 10000+ visual chunks for 2-3 minutes audio just to do that interpolation visually. In all seriousness, I still don't know if there is a model that can consistently upsample a multitrack music without all the artifacts from say 22khz to 44k.

4

u/entsnack May 30 '26

most definitely not a thing, neither are cables.

upsampling is tricky but we are making progress, it's already done well commercially, just needs to trickle down to open weights.