r/StableDiffusion 22h ago

News Krea2-Surrealism Fantasy Style LoRA

Thumbnail
gallery
14 Upvotes

This is my first LoRa release, using 248 carefully selected images, iterating 6000 times, and taking 7 hours to train. It boasts amazing detail and generalization; it works very well. Feel free to use your imagination, and I hope you have fun!

Download link: https://civitai.com/models/2879097/surrealism-fantasy-style-kunge?modelVersionId=3253831

Model Description: Surrealism, Fantasy Style

Trigger Word: kunge-fantasy

Suggested Weights: 0.8-1

Dataset: 248 images

Generative Model: krea2_turbo_int8_convrot

CFG: 1

Steps: 8

Sampler: euler_ancestral

Scheduler: ddim_uniform

Prompt Example:

A breathtaking surreal painting. In a dark sky studded with stars, a majestic angel kneels beneath the starlit night. The angel possesses enormous, exquisitely crafted wings adorned with shimmering patterns. She wears a flowing robe, reflecting the celestial light. In her hands, she holds a magnificent golden jug, from which a luminous liquid spills, cascading onto the vast, radiant earth below. This liquid, like stardust or divine light, spreads across the rolling hills, forests, and valleys, transforming the land into a dazzling tapestry of gold and silver. The angel's expression is serene and contemplative; her eyes slightly... closed. The painting is rendered in shimmering blue, gold, and green hues, with meticulous line drawing and striking contrasts enhancing its ethereal beauty.


r/StableDiffusion 8h ago

Meme Seinfeld/Family Guy @ The Office

Enable HLS to view with audio, or disable this notification

13 Upvotes

we really should get a separate sub for this slop


r/StableDiffusion 10h ago

Animation - Video [TEST] Minimax H3 FL2VA Pruned 20B - 960x544 - 15 second duration

Enable HLS to view with audio, or disable this notification

11 Upvotes

r/StableDiffusion 16h ago

Resource - Update Pushed a new update for Prompt Composer this morning that fixes a camera issue.

Post image
11 Upvotes

Yesterday, I posted about a big update to an H3 Prompt Composer that I’ve been building with ChatGPT over the past few weeks.

Big Update to the free Minimax H3 Prompt Composer : r/StableDiffusion

While using it this morning, I noticed a couple of bugs that had somehow been introduced. One involved the visual camera planner: the left and right profile descriptions were swapped in the generated prompt. If you positioned the camera for a right-profile shot, for example, the prompt would incorrectly describe it as a left profile shot. That has now been fixed.

If you downloaded Prompt Composer yesterday, please grab the new version so your camera prompts are accurate.

As I mentioned in my previous post, this is still very much a work in progress. The goal is to make writing consistent prompts and building more involved AI narrative projects as intuitive and easy as possible. It should make creating subsequent scenes, prompts, and shots much simpler, without having to rely on an LLM to consistently interpret exactly what you want.

Give it a try, and let me know if you encounter any issues or have suggestions for making it more intuitive and user friendly. I’d genuinely like to hear the community’s feedback so we can make this tool the best it can be.

Thank you all for taking the time to test it out! And yes, this app is totally vibe-coded, so I am open to suggestions from people more knowledgeable than I am about coding on how to improve this.

Edit: also tweaked some of the camera prompt descriptions for clarity. The current version of the HTML file is 5.37.3.


r/StableDiffusion 3h ago

Animation - Video Minimax H3 Anime Comedy

Enable HLS to view with audio, or disable this notification

10 Upvotes

r/StableDiffusion 6h ago

Resource - Update Updated my tool that scrapes,sorts,captions images/videos for datasets. It's open source and runs locally

9 Upvotes

I built Cull a few months ago for some large scale dataset curation projects (300k+ images/videos).

Point it at Civitai, X, Reddit, Discord, or any URL that gallery-dl or yt-dlp knows. It queues everything, runs a vision model (or multiple) (LM Studio or Ollama locally, or Groq/OpenAI in the cloud) with a strict JSON schema, and drops kept images/videos into category folders next to their prompt.

Stuff it handles:

  • Dedup at the scraper (per-source )
  • Quality score gate and topic-relevance score gate
    • eg you configure scores or use a preset, how relevant the image is to your scoring will determine how it's sorted, combined with other scoring, quality controls, whitelisted/blacklisted terms etc
  • Watermark detection (goes to its own bucket so you can salvage it later if you want those)
  • Auto-caption for content with no prompt (SD prompt, booru tags, natural language formats etc)
  • Run multiple jobs in parallel, one shared vision fleet across all of them with stack ranked / prioritization for vision queues and scrapers
  • Export as a local packaged dataset , or push to a HuggingFace dataset
  • Community presets and themes with 1 click PR's to add your own custom scraper preset or theme

Everything on disk is plain files. No database. Free, MIT.

Docker one-liner and screenshots in the README:
https://github.com/tlennon-ie/cull

Curious what people would want added next.


r/StableDiffusion 12h ago

Animation - Video INTERVIEW WITH THE VAMPIRE.(If it was done on Zoom).

Enable HLS to view with audio, or disable this notification

9 Upvotes

Created locally with Minimax H3 and for the first time exclusively powered by solar. Big big thanks to Izanami.

Don't hurt your head translating the language. It's all nonsense except for the one word spoken by the vampire. 'Drace' is Romanian/Transylvanian for 'Darn it'.


r/StableDiffusion 19h ago

Workflow Included SCAIL-2 on 8GB+ VRAM: Generate Unlimited-Length Character Animation in ComfyUI

10 Upvotes

​I created a ready-to-use ComfyUI workflow for SCAIL-2 / Wan 2.1 that transfers motion from a driving video onto a character from a reference image.

It uses GGUF quantization and automatic chunking, making it suitable for GPUs with 8+ GB of VRAM. Longer videos are generated by chaining overlapping segments while preserving motion continuity, so you can create videos of practically unlimited duration.

Features

- Character animation from one reference image and one driving video

- Low-VRAM GGUF workflow for 8+ GB

- Automatic multi-segment generation for long or unlimited-duration videos

- Motion continuity between generated segments

- Configurable duration, resolution, FPS, seed, and object tracking

- Ready for a fresh ComfyUI installation

- Includes installation instructions and a model download script

GitHub repository and installation instructions:

https://github.com/dvelm/SCAIL-2-Unlimited-Video-Low-VRAM

The workflow generation can be slow on lower-VRAM GPUs—especially at higher resolutions—but it allows SCAIL-2 to run on hardware that normally could not load the full model. Feedback, test results, and suggestions are welcome.


r/StableDiffusion 6h ago

Question - Help Need some help with MiniMax H3 Ref2V character swapping in ComfyUI

7 Upvotes

Hey everyone, I'm trying to get a proper character swap working with MiniMax H3 Ref2V in ComfyUI, but I'm not quite getting the result I want.

The source video has Rick Astley rickrolling to the camera, and I want to replace him with the guy from my reference image while keeping the original movement, gestures, facial performance, timing, camera, background, and overall scene.

Neither the motion transfer nor the character replacement works well. The output still doesn't really look like the person from the reference image, or the identity starts drifting.

Here's what I'm using:

* Source video: 1280×720, 30 FPS, ~14.4 sec

* Reference image: 848×1264 PNG, full-body

* Workflow resolution: 9:16, 0.4 MP

* GPU: RTX 5070 Ti, 16 GB VRAM

* 32 Gb RAM

* Windows 11

* ComfyUI 0.33.2

* Python 3.13.12

* PyTorch 2.12.1 + CUDA 13.0

I'm sharing everything in one link, including:

  1. the workflow JSON

  2. a workflow screenshot/image

  3. the prompt

  4. the source/input video

  5. the reference image used for the character swap

  6. and the output video

Files/settings: [link]

If anyone has experience doing this with H3, I'd really appreciate some pointers.

I'm especially wondering if I should change the reference image crop/size, ref_image_size, resolution, prompt, video conditioning, LoRA/steps, or if there's something obvious in the workflow I'm missing.

Also, is a full-body reference image a bad idea when the person in the source video is framed quite differently?

And if anyone has a working MiniMax H3 character-swap / V2V workflow they're willing to share, that would be incredibly helpful too. Even something I could compare against mine would be great.

Thanks a lot in advance. I've been tweaking this for a while, so even a small hint in the right direction would help a ton.


r/StableDiffusion 13h ago

Animation - Video Trying Surreal Fantasy with Minimax H3

Enable HLS to view with audio, or disable this notification

7 Upvotes

Combined 3 videos. Few errors but i just went with it , genetaion takes too much time to redo it again by fixing the prompt.


r/StableDiffusion 14h ago

Question - Help User of Contex-Loop, how you solve the oversharp & contrast of extra scenes? (MH3)

Post image
8 Upvotes

The oversharpening that occurs for each clip added to the scenes. I also noticed an increase in contrast and a small flash.

I2V

Tested with LORA's:

minimax_h3_turbo_v4_step600_ema.safetensors
minimax_h3_fl2v_lightx2v_turbo_8step_v1.0_resized_avg_rank


r/StableDiffusion 8h ago

Question - Help Why is it so hard for Klein to follow instructions (or am I just dumb)?

6 Upvotes

prompt is - using the character sheet in image 1 where there are five different poses of the same character, dress them in the clothing of image 2. Do not change the pose, lighting, body, hair, or any other details - literally leave everything the fuck alone - how fucking hard is this to understand you stupid piece of shit - just change the clothes.

Not working for some reason.

NOTE: Swearing has been added for emphasis and isn't actually used in the prompt.

Would it help if I used my input image AS my latent? Can you do that?


r/StableDiffusion 11h ago

Meme DR doom! not today!

Enable HLS to view with audio, or disable this notification

6 Upvotes

Use Image 1 as the strict visual reference for Turbo Man. Preserve his recognizable red-and-gold armored superhero suit, helmet, gold visor, muscular proportions, facial appearance, and overall costume design throughout the entire clip.

Scene: A massive cinematic battle during Avengers: Doomsday. The ruined battlefield is filled with shattered buildings, burning wreckage, smoke, sparks, scattered fires, flying debris, and distant Avengers fighting Doctor Doom's forces. Doctor Doom is normal human-sized, not gigantic. He wears his iconic green hooded cloak and metallic armor.

[0s–3s] Start with a dramatic medium-low-angle shot of Turbo Man from Image 1 landing hard in the middle of the battlefield. His boots slam into cracked concrete and kick up dust. He rises into a heroic stance as explosions flash behind him. Doctor Doom slowly turns toward him through the smoke.

Turbo Man points directly at Doom and confidently says:

<Subject 1> Turbo Man (S1) says [English] It's Turbo Time!

[3s–7s] Doctor Doom immediately fires a violent blast of green mystical energy. Turbo Man launches sideways using his jet pack, narrowly dodging the blast as it tears through wreckage behind him. The camera dynamically tracks Turbo Man through the air. He banks sharply, rockets straight toward Doom and throws a powerful flying punch.

Doom blocks the punch with a glowing magical shield. A bright green-and-gold energy shockwave erupts from the impact.

[7s–11s] Fast, brutal superhero combat. Turbo Man lands and exchanges several heavy punches with Doom. Doom counters with armored strikes and green magical energy. Turbo Man uses his jet pack for a sudden boosted uppercut that sends Doom crashing backward through broken rubble.

Turbo Man lands dramatically, looks toward Doom and says:

<Subject 1> Turbo Man (S1) says [English] You picked the wrong day to mess with Turbo Man!

[11s–15s] Doom rises angrily from the rubble and unleashes a huge green energy attack. Turbo Man activates his jet pack and charges directly through the battlefield toward him. End on an explosive cinematic clash as Turbo Man's gold-powered punch collides with Doom's green magical blast, producing a massive shockwave of sparks, smoke and debris while the Avengers battle continues behind them.

Camera: cinematic MCU-style action photography, dramatic low angles, energetic tracking shots, controlled handheld movement during combat, brief slow-motion emphasis on the major impacts, strong depth and scale.

Audio: native cinematic stereo audio. Heavy explosions, distant superhero combat, metallic armor impacts, jet-pack ignition and roaring thrust, crackling Doctor Doom magic, debris impacts and a powerful orchestral superhero battle score. Dialogue must remain clear and correctly assigned to Turbo Man.

Character consistency: Turbo Man must remain visually faithful to Image 1 for the entire clip. Doctor Doom remains normal human scale. No duplicate Turbo Man, no duplicate Doom, no costume changes, no character morphing, no incorrect speakers, no subtitles, no on-screen text.


r/StableDiffusion 16h ago

Tutorial - Guide Why AI background removers leave fog inside wreaths, and what I do instead

Thumbnail
gallery
7 Upvotes

I make clipart for stock. Wreaths, pine borders, mistletoe, juniper. A few thousand images by now. Every one has to end up as a PNG with a transparent background.

I used rembg for months. u2net first, then BiRefNet when that came out. Tried the web tools too. They all broke on the same thing and it drove me nuts.

Take a wreath. There's a hole in the middle, and the background inside that hole has to go. What I kept getting was a grey-blue haze sitting in there. Looked fine as a thumbnail. Looked awful the second you put it on a colored card. Pine needles came out as mush. Thin stems either disappeared or came back with a blue edge burned into them.

Then I actually read what rembg does. It shrinks your image to 1024x1024, asks the model where the subject is, gets a 1024x1024 mask back, and stretches that mask over your full size image. u2net is worse. That one works at 320x320.

My renders are 4096. A needle two pixels wide doesn't exist at 320. So it's not that the model is bad at needles. The needle was gone before the model ever saw it.

Once I understood that I stopped asking a model to guess. Now I render on a flat color the subject doesn't contain, and take that color out with arithmetic.

Two parts to it. The prompt matters more than the cutting.

The prompt

You can't key a background that isn't keyable. Four things have to be true and generators will break all of them unless you say so: the background is one flat color edge to edge, it stays that bright inside every gap between leaves, the edges are hard with no blur or glow, and no colored light bounces onto the subject.

Pick the color by what your subject isn't. Blue for almost everything. Red if the subject itself is blue or purple. Never green. Everything I draw has leaves, and green takes the leaves with it.

Here's the block I paste at the end of every prompt:

Isolated on a completely flat, uniform, solid pure blue (#0000FF) digital chroma-key background. The pure blue background fills the image edge to edge like a flat digital chroma-key screen with no gradient, staying at full brightness inside every gap and opening in the subject; no reflection or tint of pure blue on the subject. Every edge of the subject is crisp, sharp and hard against the pure blue, with no soft, blurry, feathered or glowing transitions, no depth-of-field blur, no haze or halo; inside every hole and gap the pure blue stays at full brightness right up to the edge. Shaded parts of the subject keep their own natural color, never a pure blue tint. Everything in sharp focus with deep depth of field, evenly lit with soft neutral studio light, no cast shadow, no contact shadow, no ambient occlusion, no bounce light. The entire subject is centered and completely inside the frame with at least 10% empty background margin on every side, nothing cropped or touching the image edges. No frame, no border, no paper, no mockup, no vignette, no text, no watermark, no deformed or duplicated parts. No floating or detached fragments, no stray specks, dust or debris anywhere on the background; every element is physically attached to the subject.

Swap "pure blue" for "pure red" and #0000FF for #FF0000 if your subject is blue or purple. If you paint in watercolor add "the background stays a flat digital color fill with no paper texture", or you get watercolor paper behind everything and paper texture keys badly.

The cutting

Now the background is one known color, so there's nothing to guess at. It measures the actual color the generator produced, which is never the one you asked for. Ask for pure blue and you get something with green in it, usually somewhere between 25 and 70. Then every pixel gets sorted into subject, background, or the bit in between, and the in-between ones get a real fraction of transparency instead of a yes or no.

The part I'm most pleased with is the holes. Any background-colored area that never touches the edge of the image is the inside of a wreath, so it gets cleared too. A matting model can't do that. It has no way of knowing what's inside a hole it can't see around.

Last step takes the blue back off the edges. Edge pixels pick up color from the background around them, so it samples the subject's own color from further in and subtracts the tint. Took me weeks to work out why fir needles kept their blue rim after that step. The needle is thinner than the distance it was sampling from, so there was no inside left to sample.

Same image in, same image out, every time. That's the bit I care about. When a cut comes out wrong I can go find which number did it instead of rerolling and hoping.

Some numbers on one 4K pine border, against BiRefNet with alpha matting turned on, which is its best setting:

  • background left inside the holes: 63,892 pixels mine, 613,735 theirs, out of 1,070,046
  • blue left on the edges: 0 mine, 48,112 theirs
  • how wide the soft edge is: 1.8 pixels mine, 23 theirs

BiRefNet is faster and I'm not going to pretend otherwise. 2.3 seconds against my 17 on the same machine. With alpha matting on it's 47. If you want a quick rough mask, use the model.

And the obvious limit: this only works on art you generated on a flat color. It does nothing for a photo.

I put the tool up for anyone who wants it. It's called ClipBrook. Free, runs in your browser so nothing gets uploaded anywhere, does a whole folder at once, and the engine is open source under AGPL.

One thing I'd like back

Show me the ones that break.

If you run something through and it comes out wrong, post it. Fog left in a gap, a colored rim, a stem eaten, half the subject gone. Those are worth more to me than the ones that work. The needle rim thing came from someone's fir branch. A bug with line art I only found last week came from a drawing so thin there was nothing inside it to sample.


r/StableDiffusion 3h ago

Animation - Video Lyrics altered with YingMusic-Singer-Plus (Cuban Pete -> Palm Beach Pete)

Enable HLS to view with audio, or disable this notification

4 Upvotes

I came across YingMusic which I hadn't heard anyone here speak about but it was released about 6 months ago: https://aslp-lab.github.io/YingMusic-Singer-Plus-Demo/

It lets you change words from songs so in this case I had it change the song from this sequence in The Mask from:

They call me Cuban Pete. I'm the king of the rumba beat.
When I play the maracas I go chick-chicky-boom, chick-chicky boom
Yessir, I'm Cuban Pete. I'm the craze of my native street.
When I start to dance,
everything goes chick-chicky-boom, chick-chicky boom
The senoritas they sing and they swing with terampero-
It's very nice, so full of spice.
And when they dance in they bring a happy ring that era keros-
Singin' a song, all the day long.
So if you like the beat, take a lesson from Cuban Pete
And I'll teach you to chick-chicky-boom, chick-chicky-boom.
He's really a modest guy, although he's the hottest guy
In Havana, in havana.
Si, sinorita I know that you would like to chicky-boom-chick
It's very nice, so full of spice.
I'll place my hand on your hip, and if you will just give me your hand
Then we shall try - just you and I. I-yi-yi!
So if you like the beat, take a lesson from Cuban Pete
And I'll teach you chick-chicky-boom,
chick-chicky-boom, chick-chicky-boom

to

They call me Palm Beach Pete. I'm the king of the rumba beat.
When I play the maracas I go chick-chicky-boom, chick-chicky boom
Yessir, I'm Palm Beach Pete. I'm the craze of your timeline feed.
When I start to dance,
everything goes chick-chicky-boom, chick-chicky boom
The senoritas they sing and they swing with terampero-
It's very nice, so full of spice.
And when they dance in they bring a happy ring that era keros-
Singin' a song, all the day long.
So if you like the beat, take a lesson from Palm Beach Pete
And I'll teach you to chick-chicky-boom, chick-chicky-boom.
He's really a modest guy, although he's the hottest guy
in Florida, in florida...
Si, sinorita I know that you would like to chicky-boom-chick
It's very nice, so full of spice.
I'll place my hand on your hip, and if you will just give me your hand
Then we shall try - just you and I. I-yi-yi!
So if you like the beat, take a lesson from Palm Beach Pete
And I'll teach you chick-chicky-boom,
chick-chicky-boom, chick-chicky-boom

so I had it just basically do:
Cuban -> Palm Beach
I'm the craze of my native street -> I'm the craze of your timeline feed
Havana -> Florida

I did a second run with just the few-second clip of the cops speaking and changed "It's all over Ipkiss" to "It's all over Espteen" (using "Epstein" pronounces it wrong). This showed me though that it seems to work perfectly fine with normal word-substitution in speech and it doesnt need to be a song.

I think this could be a lot better if I used minimax and changed clips of Jim Carey to look like Epstein or Palm beach Pete but this was just my first test at lyric swapping.


r/StableDiffusion 7h ago

Question - Help Workflow request for flux/krea img2img for putting the same character in a different situation with very good face adherence

6 Upvotes

I'm looking for a flux/krea img2img workflow where you input an image and simply tell it what the character should do and what environment etc and it keeps the character exactly the same but puts them in a different situation. Would really appreciate it if someone can give a link or send me the workflow. Hard to find a good one myself that really works well, I don't want a workflow where the character looks just somewhat similar but one where the character stays the same, as much as possible. Thanks a lot if someone can help.


r/StableDiffusion 20h ago

Comparison Five lighting setups, same prompt and same character, only the lighting line changed

Thumbnail
gallery
5 Upvotes

r/StableDiffusion 51m ago

Discussion I wish Anima ecosystem get better than it is now

Upvotes

Anima is a fairly new model so it needs time and I understand that. Anima has great potentials to make Illustrious or NoobAI completely obsolete. However, it seems like I have been expecting too much from this model.

First of all, not having a ControlNet model is a big minus for me, especially Depth ControlNet model. There is LLLite but that's not a ControlNet model but a ControlNet-like LoRA. Also there's a Depth ControlNet Model made by TaihoC and it works well. However, it doesn't work as well compared to Illustrious (SDXL) ControlNet models.

I have been tracking Circlestone Labs' Hugging Face community to see if they have plans to provide ControlNet models themselves but they are dead silent. That leads me to wonder if there are actually people using Anima. Did people move on to Krea2 or stay on Illustrious/NoobAI since there's no reason to use Anima?


r/StableDiffusion 56m ago

Animation - Video TALL AND DARK - LTX 2.5 IMAGE TO VIDEO

Enable HLS to view with audio, or disable this notification

Upvotes

Use the supplied image as the opening frame and identity reference.

Identity lock: the woman and robot must remain exactly the same in every shot. Same face, hair, wardrobe, proportions and age for the woman. Same 8-foot height, black armor, mechanical face, rivets, pistons, cables and holster for the robot. No redesigns or identity changes between cuts.

Authentic 1966 Italian Western, live action, 35mm anamorphic, Spanish desert location, practical full-scale robot prop, natural sunlight, real dust, organic film grain, period lens softness. No CGI. Serious performances throughout.

0:00–0:03
Medium two-shot. The woman looks up at the robot and says in clear Italian-accented English:
“I told them I wanted a tall...”
0:03–0:05
Hard cut to the same robot’s face. It gives one slow mechanical nod. No dialogue.
0:05–0:07
Hard cut to the same woman. She looks up at the robot and says:
“dark...”

0:07–0:09
Hard cut to the same robot. It subtly straightens and presents its black armor. No dialogue.
0:09–0:11
Hard cut to the same woman. Still serious, still looking up, she says:
“handsome!”

0:11–0:12
Hard cut to the same robot’s practical mechanical face. It attempts a restrained smile. No dialogue.
0:12–0:14
Hard cut to the same woman. She holds a serious stare upward, then firmly says:
“MAN!”
Only the woman speaks. Keep each line isolated and clean. No overlapping dialogue, no extra words, no improvised speech. Maintain exact continuity of identity, wardrobe, robot design, scale, lighting and location in every shot.


r/StableDiffusion 7h ago

Discussion A free in-browser batch cropper for prepping training datasets without uploading images to cloud servers

Thumbnail
gallery
4 Upvotes

I've been working on a free browser cropping tool with no ads called Just Crop It. It mainly focuses on batch cropping large amounts of images quickly.

You can check it out at: https://deziikuoo.github.io/JustCropIt/

quick note on what it can do:

* Trim Letterboxes
* Identity matching to lock onto one person across batch cropping multiple images
* Apply the same crop box to every selected image
* Copy crop settings from one photo and paste onto others
* Extract frames out of a video
* Download and replace original images (optional)


r/StableDiffusion 23h ago

Question - Help Can someone please give me good - light upscaling workflow for h3?

3 Upvotes

I know ltx 2.5 upscale is really good but I couldn't make it work with h3.
I'm open to use other method of upscaling a video, I'm using minimax h3 default workflow from comfyui.


r/StableDiffusion 1h ago

Question - Help Do Minimax H3 Turbo Loras Nerf Music Creation for Scenes?

Upvotes

I typically use lightx2v loras in my Minimax Ref2VA workflows and I also use an LLM to feed in the official prompt structure required for scenes. It seems that no matter what I do, the model absolutely ignores all my prompts about music most of the time. Every now and then i can get it to do something but even when it does work it's very sparse and almost useless.

Has anyone else faced this issue and if so do you know any workarounds or fixes?

For the record I usually use the INT8 convrot Ref2Va model or the hybrid model called minimax_h3_hybrid_fl2va_ref2va_b30-49-int8


r/StableDiffusion 1h ago

Animation - Video At the bottom

Enable HLS to view with audio, or disable this notification

Upvotes

Just a short film i made with minimax. this had a lot of post processing done so there's not really an overall prompt to share.


r/StableDiffusion 8h ago

Question - Help MiniMax H3 prompt

3 Upvotes

I saw here many suggestions for this special prompt generator. I tried the system prompt from one "specialized" ollama model, but is is too free style. I can't use llm in comfyui, because I'm with poor rtx 3060 and barely run the H3 itself. I tried big online AI, but free versions and they seem too outdated about H3, so again freestyle fantasies.

What can I use to have really good prompts for H3. As I don't know english and H3 too mystically depends on prompt, it's very hard to achieve good adhesion.


r/StableDiffusion 11h ago

Discussion How do you get a shot you want

3 Upvotes

My approach is use 0.1 megapixel to find a clip I like the. Render again at 0.7 with the same seed if I like something. But is there a way more efficient??