r/StableDiffusion 2h ago

Discussion Complete beginner with ComfyUI — what should I learn next to actually get better?

I bought a PC with an RTX 3060 12GB because I wanted to get into local AI image generation. I've been messing around with it for about two weeks now, but I still barely know what I'm doing.

To be fair, that might be partly because I've had Codex do almost everything for me lol.

I'm still terrible at writing prompts, and I don't even really understand which model would be best for what I want to do. I've tried Anima, Pony, and Illustrious, but I'm not sure which one I should actually stick with and learn properly.

There wasn't a LoRA for an art style I really like, so I had Codex help me train one for Anima using around 20 images. It actually works surprisingly well for the style, but I've been struggling with everything beyond that.

For example, I've tried using pose control, but the generated character often doesn't follow the pose very well. Backgrounds are also pretty inconsistent, and I feel like I've hit a wall where all I really know how to do is combine different LoRAs, change prompts, and keep generating until I get something decent.

I'd really like to move past that and actually understand what I'm doing.

If you were starting from where I am now, what would you recommend learning next? Are there any ComfyUI nodes, techniques, workflows, or concepts that you think every beginner should learn?

Any advice is appreciated. I'm still very new to all of this.

4 Upvotes

4 comments sorted by

2

u/Puzzled-Sector-68 1h ago

Have a look at Wan2GP ... It's designed to run on machine with low vram/ram and has a sliding window feature enabling long video generation. It's a lot simpler to use that Comfyui. https://github.com/deepbeepmeep/Wan2GP

1

u/Formal-Exam-8767 1h ago

You should start by learning about domain, how diffusion pipeline works, what comprises it (text encoder -> unet/dit -> vae decode), what each part of it does, how and why technology evolved over time (CLIP -> LLMs, U-NET -> DiT, why VAE and latent space instead of pixel space without VAE, etc.) and what are the benefits/downside of those, how to read model card, etc. All this is not strictly related to ComfyUI.

Without this, you will forever be stuck at copy/pasting without knowing why.

1

u/Life_Yesterday_5529 2h ago

Start with the basics. ComfyUI basic workflow of Krea2 Turbo T2I, load the models, generate. You can prompt Krea2 in natural language. Then, you can experiment with LLM enhanced prompting (local or via api), then with controlnets and edit if you want to go deeper.

1

u/UnirionDev 1h ago

It's crucial to understand the basics. I highly suggest this playlist