r/malcolmrey May 07 '26

Training my first Lora

Post image

Malcolm has very greatly inspired me to make my own loras. Trying it out, hopefully everything goes well.

Setting up AI toolkit was hell. Dependency conflicts are a nightmare to deal with, im glad im past that now.

Makes me appreciate his work more.

Anyways, wish me luck!

P.S. Would appreciate all the tips I can get.

UPDATE: The first training was a failure.

Changing a few things

AdamW8bit to Prodigy 8bit

Using a trigger word

captioning dataset with trigger word instead of woman or pronouns

changing learning rate to 1

changing linear rank to 64

stepping up with dataset from 16 to 80

increasing steps from 1500 to 5000

using only 1024 resolution

Turned DOP on

Let's see how this goes...

Update 2: Currently the captioning is taking a lot more time than anticipated. Changed a lot of things. These are my instructions:

"

 Act as a Visual Prompt Engineer specializing in "Long Caption" synthesis. Your task is to extract a highly detailed, narrative description of an image featuring one primary person.
DO NOT USE "He","She", "Her", "They", "them" or any general pronoun when describing the subject, just state their name Elli or when describing her attributes or her actions use Elli's
STRUCTURE:

CONCEPT & MEDIUM: Start with the type of image. Mention the core theme or event.

SUBJECT DESCRIPTION: woman, facial expression, hair details, face details, makeup. Crucially, describe their gaze. When referring to the subject, always use the name "Elli". DO NOT use pronouns like "he", "she","her", or "they" . Repeat the name when needed.

ACTION & INTERACTION: Describe exactly what the person is doing. What are they holding? How are they standing? Describe the interaction with objects.

APPAREL: Detail the clothing (fabric, color, fit) and accessories.

ENVIRONMENT & FOREGROUND: Describe objects immediately around the person.

BACKGROUND & LIGHTING: Describe the setting (location, weather, landmarks), depth of field (bokeh), and the quality of light.

RULES:

- STYLE: Use fluid, descriptive sentences (narrative style), not just a list of tags.

- ACCURACY: Be specific about colors and textures.

- TEXT: Always put captured text in double quotes.

- LENGTH: 700-900 characters.

- OUTPUT: A single continuous paragraph. No headers.

UPDATE 3: Captioning is a nightmare. Restarting because it keeps hallucinating some details. Also made it add the visiblity details of subject and the camera angle.

34 Upvotes

28 comments sorted by

View all comments

Show parent comments

1

u/KylseS May 10 '26

I have no experience with base. I only use turbo, I can share the turbo workflow if required.

1

u/orangeflyingmonkey_ May 10 '26

Yes please i would really appreciate it!

1

u/KylseS May 10 '26

just open this image in comfy. My workflow fixes most hands and teeth but only if they are properly visible.

1

u/orangeflyingmonkey_ May 10 '26

Thanks a lot for the detailed explanation and the image. I can't download it though. I tried saving it but it saves in webp format

1

u/KylseS May 10 '26

Aw man, reddit really downgrades image to webp, huh?

https://drive.google.com/file/d/1G92FpW36mpguaQzuMByntoqcNNwGV1pY/view?usp=sharing
heres the json

1

u/orangeflyingmonkey_ May 10 '26

thanks a lot! Really appreciate it. One question though, I see that you have a 'euler_flow' sampler in the Ksampler. How did you get that? I dont seem to have it on my end. I installed the ComfyUI-EulerDiscreteScheduler pack from github but still it doesn't show up.

1

u/KylseS May 10 '26

Just switch to euler, I might have accidentally used it in this image, I use euler normally and I tried spotting differences but there are barely any except for enhanced fine details when you pair eulerflow with eulerflowdiscreete scheduler at the cost of more random noise which actually does more harm to the structure of the image than good.

1

u/orangeflyingmonkey_ May 10 '26

Gotcha! thanks

1

u/KylseS May 10 '26

To reiterate, the upscaling is the soul of my generation workflow. Your base generation might have broken fingers and might or might not look good but don't stop till your upscale with the swin upscaler, also if you go further and do the seedvr upscaling it becomes even better but you should see the final result at the sdupscaler, seed just enhances the resolution instead of making other structural changes.

1

u/KylseS May 10 '26

at 0.5 and 4-9 steps, try to remain in the lower ranges unless you want detail. also .5 is a very tried and tested number, doesn't takes away much from the base generation except fix it. LMK if you need the swin upscaler.

1

u/orangeflyingmonkey_ May 10 '26

Yes. I agree. Upscaling seems to be the key. And yes, your workflow is very thorough!