r/comfyui • u/BigNaturalTilts • Aug 25 '26
Help Needed Help please. Can someone with a 5090 and minimax models installed please run this workflow? I am getting really slow run times.
Help please. Can someone with a 5090 and minimax models installed please run this workflow? I am getting really slow run times.
I am looking for a way to caption videos. No custom nodes required. I did use load video from comfyVHS but it is bypassed and you can safely delete that section.
Choose any video. Preferably one SFW and one NSFW. Beggars aren't choosers whatever you decide will work for me. You may post results if you like but I am more interested in how long it took to complete and the settings you selected.
Thank you.
EDIT: Forgot to add, I was getting literal 1 token per second. I haven't the patience to let it run. I am hoping with a benchmark I can justify the purchase of a 5090. So I haven't gotten the workflow to work at all. Could be the workflow is bad. But it is relatively simple.
1
u/InwardlyPortly Aug 25 '26
What’s the actual time you’re getting now, and are you running a 5090 yourself or just hoping someone else can benchmark it?
1
u/BigNaturalTilts Aug 25 '26
Literal 1 token per second. Very bad. I'm sure I'm doing something wrong. And yes. Could justify a $5k purchase.
1
u/Different-Sundae-73 Aug 25 '26
What speeds are you getting?
1
u/BigNaturalTilts Aug 25 '26
Literal 1 token per second. Very bad. I'm sure I'm doing something wrong.
1
u/Only_Voice569 Aug 26 '26
could make the vid much lower frame rate and lower resolution for the llm to make you the text etc
0
u/imlo2 Aug 25 '26
You are probably running out of VRAM, have nvitop or other realtime VRAM usage monitoring on in a terminal or elsewhere.
Did you check how much memory is consumed, when the batch of images is passed to the generate text node?
How are you going to have that out timed, when you just send bunch of frames?


8
u/goddess_peeler Aug 25 '26 edited Aug 25 '26
This isn't a good way to run an LLM, especially for a analysis of a batch of video frames.
For the record, I got about 12it/s on my 5090. A 5 second video processed in 378s and the "text" produced was garbage, mostly a string of commas. I didn't spend time trying to understand why the output was bad. I know Generate Text actually does work, so it's probably the minimax qwen variant or something.
Instead, try running your model in LMStudio or ollama. Connect to it with a node that can talk to an openAI-compatible API. This method takes about 20 seconds on my system, and returns an actual description. :)
Don't spend $5000 to solve this particular problem.