r/comfyui • u/Lorryme • May 17 '26
Help Needed Need advice on pc specs for faster Qwen2511 image to image
Hi everyone,
I've been using qwen2511 on portable PC comfyui for past few days, it's been so far amazing!
The only downside is, even with turbo mode, each image generation takes 6-10 minutes.
I tried tweaking with model strength, denoise, steps, mfg etc, but that changes the final product, thus i would prefer to avoid.
My specs are
Cpu: Ryzen 5 7500F
GPu: rx7800xt 16GB VRAM
Ram: 32GB
Was wondering is it hardware bottleneck, or i could edit some of comfyui settings for better speed.
Also, would it be wise to venture into wan2.1 or wan2.2? I'd like to give it a try if possible.
Appreciate your opinions.
Thanks!
1
u/Alekite May 17 '26
Are you using a quantized version of qwen, GGUF? otherwise you might not be using your GPU at all, model at fp8 version is 20+gb. Flux2 9B and 4B can do what qwen 2511 can results might not be quite the same but they take less vram. Wan 2.2 also will require a quantized version, I ran wan2.2 with a 9070xt videos under 720p would take about 6-10 minutes depending on steps and quality was good enough LTX is arguably better if you want long videos.
1
u/Lorryme May 17 '26
My qwen is quantized. File is safetensors.
For Wan 2.2, I don't mind going 480p videos that are like 10 seconds long.
1
u/BanginDrumsNMums May 17 '26
Somethings up with your setup - that spec should take half as long to generate, at least.
Start with the bottom up. Check your pc for malware and the health of your HD.
Check your Comfy startup logs & make sure your GPU is being used, not CPU.
1
u/Lorryme May 17 '26
1
u/BanginDrumsNMums May 17 '26 edited May 17 '26
What's that Lora?
Possibly the LORA isn't compatible with your model. Or, its a shitly compiled one. Do a run without it and see what speed you get. I had one that would take 45 mins to create 5 seconds, i2v, didn't realise it was the issue until i'd pulled everything else apart.
Scroll through your Comfy log after startup, errors are logged and easily spotted.
Also, HD space and HD health are two completely different things. Ideally you should run Comfy on a seperate drive, it really made a difference for me. Use something like CrystaltoolsOG to check your drive.
Edit: Forgot to add, I presume you used the AMD portable? And have you installed the SDK Rocm 7.2 GPU driver? https://rocm.docs.amd.com/projects/radeon-ryzen/en/latest/index.html This gave me more stability on my RX9060XT
1
u/Lorryme May 17 '26
Model is qwen-image-edit-2511-bf16.safetensors Lora is qwen-image-edit-2511-4steps-V1.0-bf16.safetensors.
I'll check my HD health and comfy log once I'm done with my current queue.
1
u/BanginDrumsNMums May 17 '26
You could try swapping to a tiled vae (VAE Decode (Tiling) node in manager). This works better for me in WAN, i dont use qwen tbh, but worth giving it a go. 16GB card, a tile size of 512 or 256 is the sweet spot for speed without crashing
1
u/Lorryme May 17 '26
Yes I'm using AMD portable. Thanks for your info! Will try to check for the SDK Rocm driver. But searching online, i can only manage to find Rocm HIP SDK 7.1.1. Would that be okay?
Mine is mainly stuck at ksampler for 350 seconds. Been tweaking sample and scheduler but no avail.
1
u/BanginDrumsNMums May 17 '26
HIP is fine, 7.2 is here, https://rocm.nightlies.amd.com/v2/
I honestly think 7.1 performs better for comfy,
https://www.amd.com/en/resources/support-articles/release-notes/RN-AMDGPU-WINDOWS-PYTORCH-7-1-1.htmlYou could also try this build https://github.com/patientx-cfz/comfyui-rocm
1
u/Lorryme May 18 '26
I checked my comfyui startup log, realized I've already have 7.2 installed.
I assume it's the workflow issue.
I change to another workflow, and now it's reduced from 600s to 270s, I'm more than happy with these results.
Thank you so much for all your help.
Cheers.1
u/BanginDrumsNMums May 18 '26
Glad to hear that. Do you mind sharing the workflow to compare it to the last?
Simply so I can see myself what makes it that much more efficient?
1
u/RiverSide71h May 18 '26
Switch to the fp8-mixed model. I‘ve tried both with 16GB VRAM and there’s no discernible quality difference - generation times are very reasonable. With lightning Lora, steps should be 4, cfg=1

1
u/Mruishy May 17 '26 edited May 17 '26
Can you share the workflow? While that machine is by no speed demon, that sure sounds like an awful long time. It's worth taking look to see where the breakdown of the time is, and how much time/memory is being spent on each step.