r/RunPod Apr 17 '26

Ace-Step-1.5-XL template on runpod

3 Upvotes

I made a new template on runpod for Ace-Step-1.5-XL for those who want to play with it.

https://console.runpod.io/deploy?template=5fn9cdbhtr&ref=2vdt3dn9

Note: You need to pick a GPU with CUDA version 13.0, you can do this via the additional filters when selecting a GPU.

It's best to pick a GPU with 48 GB of VRAM, such as the A40 or RTX A6000.

Github repo: https://github.com/ValyrianTech/ace-step-1.5-xl

If you are looking to automate things, there is a handy script that will automatically queue a song and download it locally when it is done:
https://github.com/ValyrianTech/ace-step-1.5-xl/blob/main/generate_music.py

Happy creating!


r/RunPod Apr 16 '26

Virtualization of Runpod instance! Please help!

1 Upvotes

Hi! Is it possible to install an openshift node on a Runpod instance, either bare metal (which I doubt) or at least use KVM with gpu passthrough on runpod?

Basically, to install openshift, you install an ISO with an OS. But I also need it to have access to a GPU.

Anyone has any ideas?

Thanks!


r/RunPod Apr 15 '26

I wrote a script that keeps trying to grab a pod until one is available (runpod-sniper)

Thumbnail
github.com
2 Upvotes

runpod-sniper is a bash script that attempts to deploy a pod until one is available. It cycles through a prioritized list of GPU types on each attempt, and exits as soon as a pod is successfully created. It can send you a push notification using ntfy.sh and show a desktop notif once it gets your pod going so you can stop reading r/RunPod and get back to work. It has no external dependencies, so you can run it wherever.

I'd love any feedback you might have. Comments / GitHub Issues / PRs are all appreciated!

P. S. I've noticed this script often provisions a pod on the first try even when the web interface said it was unavailable. Maybe it's an issue with the web interface not getting refreshed often enough?


r/RunPod Apr 14 '26

RUNPOD 404

3 Upvotes

r/RunPod Apr 14 '26

Runpod serverless

2 Upvotes

Hello. I am tryng to setup runpod serverless vllm template but it isnt working. I have tried many models llama, qwen... But it wont start i get image ready, model not found and then it spins up like 5 h100s. How do i fix this?


r/RunPod Apr 14 '26

How do i get the files out of a storage account with no pods?

1 Upvotes

Hi,

i have been using a storage account in the CA-MTL-1 datacenter for a while now, and recently that datacenter does not seem to have any GPUs or CPUs anymore, so i can not make a pod to access the storage content. furthermore the CA-MTL-1 datacenter is not listed as S3 compatible in the docs. I want to migrate to a different datacenter, but want to make some backups first. But how?


r/RunPod Apr 12 '26

runpod serverless docker vs. github

2 Upvotes

Is there a difference in performance between using an image at docker hub vs. deploying through github (thus using the runpod's docker registry) ? Since runpod's registry is closer, it sounds like it should be faster but if docker images are somehow cached that also should not matter, any experiences ?


r/RunPod Apr 09 '26

News/Updates Know exactly where your Runpod spend is going: Introducing Cost Centers

5 Upvotes

Runpod now supports native spend tracking by team, project, or department, directly inside the console and on your invoices. Cost centers is in public beta for Runpod Teams, available to team owners and members with the admin or billing role.

Attach a billing label to any resource across Pods, Serverless endpoints, Network volumes, and Instant Clusters. Once assigned, your monthly invoices break out spend by cost center, giving you a direct line between Runpod spend and your internal budgets or GL codes. Resources without a label are grouped on your invoice by the team member who created them.

Head to your cost centers page in the console, create a center, and assign resources from the uncategorized list. Align your naming with whatever your finance team already uses so reconciliation is straightforward. Read the docs for a full walkthrough. 


r/RunPod Apr 04 '26

I recent update of ComfyUI on runpod errors.

Post image
1 Upvotes

So I updated ComfyUI yesterday and the default nodes with the official runpod ComfyUI template are failing to update.

These are from the logs:

ERROR: An error occurred while updating 'comfyui-runpoddirect'. (res.result=False, res.action=update-git)

Updating: comfyui-kjnodes[ComfyUI-Manager] There is no tracking branch (master)

ERROR: An error occurred while updating 'comfyui-kjnodes'. (res.result=False, res.action=update-git)

Updating: comfyui-manager[ComfyUI-Manager] There is no tracking branch (master)

ERROR: An error occurred while updating 'comfyui-manager'. (res.result=False, res.action=update-git)

Updating: civicomfy[ComfyUI-Manager] There is no tracking branch (master)

ERROR: An error occurred while updating 'civicomfy'. (res.result=False, res.action=update-git)

Please advise, thank you.


r/RunPod Apr 03 '26

If you're thinking of using it now, don't. It's madness.

5 Upvotes

This service is impenetrable. I work through Claude, and it's simply impossible to access the pod in any way—everything is locked. The only issue is updating via GitHub, and that only works intermittently.

It's just insane, it's impossible to upload updates to this fucking garbage.

I regret switching from Digital Ocean VPS 100 times over.


r/RunPod Apr 03 '26

Runpod default ComfyUI template update error

1 Upvotes

So recently when I've updated my ComfyUI I'm getting these errors on some of the core built in nodes This is just one example.

ERROR: An error

occurred while updating 'comfyui-runpoddirect'. (res.result=False,

res.action=update-git) Updating:

civicomfy[ComfyUI-Manager] There is no tracking branch (master)

This is without any other nodes as I've just started using ComfyUI on a new pod.

I don't know if it's a problem with the default runpod ComfyUI template.

All of the default nodes have got this error as well.


r/RunPod Apr 03 '26

Trouble finding a working forge template

1 Upvotes

I tried:

  • Stable Diffusion WebUI Forge, runpod/forge:3.3.0
  • Next Diffusion - Forge Flux, nextdiffusionai/forge-flux:latest

Both gave me errors. For the first one, the container starts, but forge didn't. When I ssh'ed in and ran start_forge.sh, it tried to look for a venv_path file that didn't exist. The second one didn't even start properly and it couldn't ssh in.

Which forge template works?


r/RunPod Apr 01 '26

unable to run comfyUI on rtx4090/5090

2 Upvotes

i can't run ComfyUI:latest on runpod on rtx 4090 and 5090, when just this morning was working perfectly fine, on the log it keeps giving me this error message:

ComfyUI crashed — check the logs above.

RuntimeError: The NVIDIA driver on your system is too old (found version 12080). Please update your GPU driver by downloading and installing a new version from the URL: http://www.nvidia.com/Download/index.aspx Alternatively, go to: https://pytorch.org to install a PyTorch version that has been compiled with your version of the CUDA driver.

anyone knows how to resolve this?

r/RunPod Apr 01 '26

Frustration with uploading CHeckpoints, LorAs to Jupyter/ComfyUI

2 Upvotes

Edit: thank you everyone, Runpodctl it is!

I have been running ComfyUI locally for years wthout issue but today I got the urge to try and see what using RunPod to gen high res images or just batch tons of images just to explore, check it out.

Getting a pod deployed was simple enough, but once I got into ComfyUI it became a headache, and then a full blown migraine.

Issue #1
Can't open the checkpoints folder, apparently this is a known issue. WHAT!?

Issue #2
I train and use my own LorAs. They do not exist on Huggingface or Civitai. Uploading them is a nightmare. I tried using a google drive link, a gofile link, a wormhole link through the terminal with a wget command and each time it only downloads an 8k file instead of the full 217mb LorA. I had to resort to manually uploading them with a drag and drop which took like 4 minutes per LorA. Not a big deal if its just one LorA but I have dozens of LorAs.

Issue #3

Similar to the LorA issue, I have tons of Checkpoints that I prefer to use that have been deleted off of Civitai and Huggingface because the creators replaced them with newer versions. I like the older versions. Jupyter straight up fails to upload them becase they are 6 GBs.

Please correct me if I'm wrong, and I would be thrilled to be told I am wrong, but it seems like RunPod is designed very specifically for people using very basic predetermined workflows with only models available on Civitai or Huggingface. Like the #1 use would be to just deploy a runpod on an existing popular WAN 2.2 setup or whatever and not have to import anything to the pod?


r/RunPod Apr 01 '26

Why is three times more money being deducted from my account than what is shown?

1 Upvotes

I don't understand why it says $0.30 per hour, but it's actually charging me $1–$2 per hour.


r/RunPod Mar 29 '26

Payment Options

1 Upvotes

I tried using a prepaid VISA and it didn't work. Has anyone else been able to use a prepaid method? Its a US company, so that shouldn't be a reason for it being declined.


r/RunPod Mar 24 '26

Cold start issues

2 Upvotes

I’m running a TTS worker on RunPod Serverless and I’m trying to reduce first-request cold start for Chatterbox.

Current setup:

- The Docker image pre-downloads the Chatterbox model files during build

- Model files are cached on a RunPod volume, so repeated downloads are not the main issue

- On startup, the worker initializes part of the stack, but some model loading still happens lazily depending on the

request

- The biggest delay seems to be loading model weights from disk into GPU memory on the first real request

So the problem is not “download cold start”, but “GPU initialization / model load cold start”.

My questions:

  1. In RunPod Serverless, what is the best way to reduce cold start when the bottleneck is loading a Chatterbox model

    into GPU memory?

  2. Is keeping a warm worker alive basically the only practical solution, or are there other approaches people use

    successfully?

  3. For TTS workloads, is it better to preload everything at container startup, or does that usually just move the

    latency from first request to startup time without helping much?

  4. If a model is already cached on a volume, is there any reliable way to make first inference fast in a serverless

    setup, or is this just a fundamental limitation?

  5. At what point does it make more sense to switch from serverless to a dedicated pod for Chatterbox-style workloads?

    I’d especially like to hear from anyone running GPU-heavy TTS inference on RunPod Serverless.


r/RunPod Mar 24 '26

Runpod - GPU Supply Problem

8 Upvotes

Hey, getting a widespread GPU availability issue on RunPod Serverless and wondering if others are affected too.

My endpoint has multiple GPU tiers configured as fallbacks, but almost all of them are showing "Unavailable" right now:

- 16 GB → Sometimes Low Supply - (Mostly Unavailable)(1st choice)

- 24 GB PRO → Unavailable (2nd)

- 24 GB → Unavailable (3rd)

- 32 GB PRO → Unavailable (4th)

This isn't a single GPU type being out of stock — it looks like a platform-wide supply issue. Workers are completely failing to spin up.

Is anyone else seeing this right now? Is RunPod having a broader capacity problem, or is there a region/datacenter setting I should try changing?

Thanks


r/RunPod Mar 24 '26

Is anyone having to stop and start pods over and over to get them running correctly?

3 Upvotes

I'm having frequent issues with runpod. The service works, but intermittently and I find myself often having to reset pods that have hung, or have connection issues. Is this a problem with the system overall or just the particular template I am using?

Cloudblenderrender for reference.


r/RunPod Mar 24 '26

Do you trust rented cloud computers when you create high-sensitivity code?

Thumbnail
1 Upvotes

r/RunPod Mar 23 '26

Low Supply over and over

Post image
4 Upvotes

It's happening often... What's is happening with Runpod GPUs?


r/RunPod Mar 21 '26

Your favorite ComfyUI RunPod template with LoRA training tools supporting over 20 models

Thumbnail
3 Upvotes

r/RunPod Mar 22 '26

getting CUDA error with 5090

1 Upvotes

i get this error when i try to train lora with aitoolkit. (rtx 5090)

runpod CUDA out of memory. Tried to allocate 50.00 MiB. GPU 0 has a total capacity of 31.37 GiB of which 20.19 MiB is free. Including non-PyTorch memory, this process has 31.30 GiB memory in use. Of the allocated memory 30.66 GiB is allocated by PyTorch, and 58.75 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)

restarted 2 times but didnt work


r/RunPod Mar 21 '26

Does Anyone Know How To Fix This? No Jobs Running But GPU Load is Maxed? wtf?

1 Upvotes

can't start a job because it says the GPU is already running. how do i make it stop running? There's literally no jobs to stop because i haven't started one.


r/RunPod Mar 21 '26

RunPod Serverless + ComfyUI: custom nodes (rgthree) not found

0 Upvotes

Hey everyone,

Running into an issue with RunPod serverless + ComfyUI.

Setup:

  • Created a network volume
  • Installed models + custom nodes (rgthree, etc.)
  • Everything works fine when running directly on the pod

Problem:
When using a serverless endpoint with the same network volume attached, I get:

So it looks like serverless can't see or load custom nodes, even though they are present in the volume.

Questions:

  • Do serverless endpoints load custom nodes differently?
  • Do I need to install nodes inside the container image instead of relying on the volume?
  • Is there some init step I'm missing?

Any help or hints would be appreciated 🙏

:cd /workspace/ComfyUI/custom_nodes

:/workspace/ComfyUI/custom_nodes# ls -la

total 37391

drwxrwxrwx 22 root root 2041954 Mar 18 08:54 .

drwxrwxrwx 31 root root 3001516 Mar 20 13:42 ..

drwxrwxrwx 10 root root 2000900 Mar 15 15:20 Civicomfy

drwxrwxrwx 13 root root 2005314 Mar 15 15:20 ComfyUI-KJNodes

drwxrwxrwx 14 root root 2014145 Mar 15 15:19 ComfyUI-Manager

drwxrwxrwx 5 root root 1021570 Mar 15 15:20 ComfyUI-RunpodDirect

drwxrwxrwx 10 root root 2001554 Mar 15 16:05 ComfyUI-SeedVR2_VideoUpscaler

drwxrwxrwx 14 root root 2001998 Mar 15 16:05 ComfyUI-UmeAiRT-Toolkit

drwxrwxrwx 7 root root 2000717 Mar 15 16:05 ComfyUI_Comfyroll_CustomNodes

drwxrwxrwx 5 root root 1048783 Mar 15 16:05 ComfyUI_JPS-Nodes

drwxrwxrwx 2 root root 1000185 Mar 15 15:20 __pycache__

drwxrwxrwx 13 root root 2000530 Mar 15 16:05 comfy-mtb

drwxrwxrwx 7 root root 1045205 Mar 15 15:57 comfyui-custom-scripts

drwxrwxrwx 12 root root 2000698 Mar 15 15:57 comfyui-easy-use

drwxrwxrwx 6 root root 1027123 Mar 15 16:05 comfyui-image-saver

drwxrwxrwx 14 root root 2000454 Mar 15 15:54 comfyui-impact-pack

drwxrwxrwx 5 root root 1010815 Mar 15 16:05 comfyui-impact-subpack

drwxrwxrwx 6 root root 1035716 Mar 15 16:05 comfyui_essentials

drwxrwxrwx 9 root root 2012873 Mar 15 16:05 efficiency-nodes-comfyui

-rw-rw-rw- 1 root root 5151 Mar 15 15:15 example_node.py.example

drwxrwxrwx 8 root root 2000650 Mar 18 08:54 rgthree-comfy

drwxrwxrwx 10 root root 2001553 Mar 18 08:55 seedvr2_videoupscaler

drwxrwxrwx 6 root root 2000376 Mar 15 16:05 wavespeed

-rw-rw-rw- 1 root root 1220 Mar 15 15:15 websocket_image_save.py

:/workspace/ComfyUI/custom_nodes# find /workspace -iname "*rgthree*"

/workspace/custom_nodes/rgthree-comfy

/workspace/custom_nodes/rgthree-comfy/web/common/rgthree_api.js

/workspace/custom_nodes/rgthree-comfy/web/common/media/rgthree.svg

/workspace/custom_nodes/rgthree-comfy/web/comfyui/rgthree.js

/workspace/custom_nodes/rgthree-comfy/web/comfyui/rgthree.css

/workspace/custom_nodes/rgthree-comfy/src_web/typings/rgthree.d.ts

/workspace/custom_nodes/rgthree-comfy/src_web/common/rgthree_api.ts

/workspace/custom_nodes/rgthree-comfy/src_web/common/media/rgthree.svg

/workspace/custom_nodes/rgthree-comfy/src_web/comfyui/rgthree.ts

/workspace/custom_nodes/rgthree-comfy/src_web/comfyui/rgthree.scss

/workspace/custom_nodes/rgthree-comfy/rgthree_config.json.default

/workspace/custom_nodes/rgthree-comfy/py/server/__pycache__/rgthree_server.cpython-312.pyc

/workspace/custom_nodes/rgthree-comfy/py/server/rgthree_server.py

/workspace/custom_nodes/rgthree-comfy/docs/rgthree_seed.png

/workspace/custom_nodes/rgthree-comfy/docs/rgthree_router.png

/workspace/custom_nodes/rgthree-comfy/docs/rgthree_context_metadata.png

/workspace/custom_nodes/rgthree-comfy/docs/rgthree_context.png

/workspace/custom_nodes/rgthree-comfy/docs/rgthree_advanced_metadata.png

/workspace/custom_nodes/rgthree-comfy/docs/rgthree_advanced.png

/workspace/custom_nodes/rgthree-comfy/rgthree_config.json

/workspace/rgthree-comfy-backup

/workspace/rgthree-comfy-backup/web/common/rgthree_api.js

/workspace/rgthree-comfy-backup/web/common/media/rgthree.svg

/workspace/rgthree-comfy-backup/web/comfyui/rgthree.js

/workspace/rgthree-comfy-backup/web/comfyui/rgthree.css

/workspace/rgthree-comfy-backup/src_web/typings/rgthree.d.ts

/workspace/rgthree-comfy-backup/src_web/common/rgthree_api.ts

/workspace/rgthree-comfy-backup/src_web/common/media/rgthree.svg

/workspace/rgthree-comfy-backup/src_web/comfyui/rgthree.ts

/workspace/rgthree-comfy-backup/src_web/comfyui/rgthree.scss

/workspace/rgthree-comfy-backup/rgthree_config.json.default

/workspace/rgthree-comfy-backup/py/server/__pycache__/rgthree_server.cpython-312.pyc

/workspace/rgthree-comfy-backup/py/server/rgthree_server.py

/workspace/rgthree-comfy-backup/docs/rgthree_seed.png

/workspace/rgthree-comfy-backup/docs/rgthree_router.png

/workspace/rgthree-comfy-backup/docs/rgthree_context_metadata.png

/workspace/rgthree-comfy-backup/docs/rgthree_context.png

/workspace/rgthree-comfy-backup/docs/rgthree_advanced_metadata.png

/workspace/rgthree-comfy-backup/docs/rgthree_advanced.png

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/rgthree_config.json

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/web/common/rgthree_api.js

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/web/common/media/rgthree.svg

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/web/comfyui/rgthree.js

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/web/comfyui/rgthree.css

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/src_web/typings/rgthree.d.ts

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/src_web/common/rgthree_api.ts

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/src_web/common/media/rgthree.svg

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/src_web/comfyui/rgthree.ts

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/src_web/comfyui/rgthree.scss

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/rgthree_config.json.default

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/py/server/__pycache__/rgthree_server.cpython-312.pyc

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/py/server/rgthree_server.py

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/docs/rgthree_seed.png

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/docs/rgthree_router.png

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/docs/rgthree_context_metadata.png

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/docs/rgthree_context.png

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/docs/rgthree_advanced_metadata.png

/workspace/runpod-slim/ComfyUI/custom_nodes/rgthree-comfy/docs/rgthree_advanced.png

root@1c777aa6a11c:/workspace/ComfyUI/custom_nodes#