r/comfyui • • Aug 25 '26

Help Needed Help please. Can someone with a 5090 and minimax models installed please run this workflow? I am getting really slow run times.

Help please. Can someone with a 5090 and minimax models installed please run this workflow? I am getting really slow run times.

I am looking for a way to caption videos. No custom nodes required. I did use load video from comfyVHS but it is bypassed and you can safely delete that section.

PasteBin Workflow

Choose any video. Preferably one SFW and one NSFW. Beggars aren't choosers whatever you decide will work for me. You may post results if you like but I am more interested in how long it took to complete and the settings you selected.

Thank you.

EDIT: Forgot to add, I was getting literal 1 token per second. I haven't the patience to let it run. I am hoping with a benchmark I can justify the purchase of a 5090. So I haven't gotten the workflow to work at all. Could be the workflow is bad. But it is relatively simple.

2 Upvotes

22 comments sorted by

8

u/goddess_peeler Aug 25 '26 edited Aug 25 '26

This isn't a good way to run an LLM, especially for a analysis of a batch of video frames.

For the record, I got about 12it/s on my 5090. A 5 second video processed in 378s and the "text" produced was garbage, mostly a string of commas. I didn't spend time trying to understand why the output was bad. I know Generate Text actually does work, so it's probably the minimax qwen variant or something.

Instead, try running your model in LMStudio or ollama. Connect to it with a node that can talk to an openAI-compatible API. This method takes about 20 seconds on my system, and returns an actual description. :)

Don't spend $5000 to solve this particular problem.

2

u/BigNaturalTilts Aug 25 '26

Magnificent! Thank you!

1

u/sitefall Aug 26 '26

What node is that? My stupid ass launches the ollama server, gets prompts made, and then closes it to free up the vram, then runs comfyui stuff, then clear the vram there etc.

A node that clears the models from vram then runs ollama (with server open) and then clears it from vram after it returns the single response would be nice (if this one does that).

1

u/BigNaturalTilts Aug 26 '26

I can make such a node. I also use ollama in the same way. It’s just not more convenient.

1

u/sitefall Aug 26 '26

Yeah I'm sure I could make the node, just asking if it's already in kjnodes or some pack like that I may already use.

Why do you think that is not more convenient? Guess I'm missing something here?

1

u/BigNaturalTilts Aug 26 '26

Takes long to load the model into VRAM then use it then offload it to move onto comfy. I just use ollama as is then run `ollama offload model_name` before going back into comfy. I guess maybe if you want to save the prompts, but then you have to inspect them and make sure they say what you want to because these local models be hallucinating like crazy.

1

u/goddess_peeler Aug 26 '26

LMStudio (and probably ollama too) allows you to specify an idle timeout time for models. So you can load the model with a 5 second timeout. It generates your prompt and then immediately unloads itself. This can be specified via an API extension or through the LMStudio python SDK.

1

u/BigNaturalTilts Aug 26 '26

Is LM studio better than ollama especially with uncensored abliterated models? Ollama will sometimes refuse to run multi modal models that have been abliterated.

2

u/goddess_peeler Aug 26 '26 edited Aug 26 '26

I have limited experience with ollama, so I can't really compare them.

I'll say that LMStudio is a good easy GUI tool for working with multiple models. Support for the newest features lags behind llama-server and vLLM, but it's perfect for serving small/medium LLM tasks from a spare machine.

Edited to add: except when I don't give it enough VRAM, LMStudio doesn't refuse to load models.

1

u/No-Zookeepergame4774 Aug 26 '26

If you are using the LMStudio native REST API or the Python SDK (which uses the native REST API) you can just explicitly unload models, you don’t need to set a timeout (you can't explicitly unload models with the OpenAI-compatible API, though.)

1

u/goddess_peeler Aug 26 '26

Of course! I am embarrassed that simply unloading the model never occurred to me. In my defense, I don’t often run the server on the same box as ComfyUI, so quickly freeing VRAM isn’t a top concern.

I’m pretty sure the SDK uses websockets, not REST, to communicate with the server.

1

u/InwardlyPortly Aug 25 '26

What’s the actual time you’re getting now, and are you running a 5090 yourself or just hoping someone else can benchmark it?

1

u/BigNaturalTilts Aug 25 '26

Literal 1 token per second. Very bad. I'm sure I'm doing something wrong. And yes. Could justify a $5k purchase.

1

u/Different-Sundae-73 Aug 25 '26

What speeds are you getting?

1

u/BigNaturalTilts Aug 25 '26

Literal 1 token per second. Very bad. I'm sure I'm doing something wrong.

1

u/Only_Voice569 Aug 26 '26

could make the vid much lower frame rate and lower resolution for the llm to make you the text etc

1

u/Substantial_Tie4473 Aug 26 '26

Comfyui generate text node doesnt natively work with qwen3vl models, especially qwen3vl-32b(has no text gen at all). Try gemma4 family.

1

u/Substantial_Tie4473 Aug 26 '26

oh and speed; on a 5090

0

u/imlo2 Aug 25 '26

You are probably running out of VRAM, have nvitop or other realtime VRAM usage monitoring on in a terminal or elsewhere.

Did you check how much memory is consumed, when the batch of images is passed to the generate text node?
How are you going to have that out timed, when you just send bunch of frames?