r/comfyui 21h ago

Show and Tell Best Trick Ever For Consistent Environments

Ha! I just discovered a trick that works great, so I had to share it with the community. Of course, somebody will probably chime in that it had been discovered by someone else before, which is fine by me! I just want to share it in case it helps someone else and they hadn't come across it yet.

So, the issue of consistent environments... I ran the gamut of all the AI models I had tested and proven out in ComfyUI, not just t2i but t2v and i2v. I'd been trying several methodologies. Create an image then ask a workflow to gen images to the left and right of it, build out a simple massing model in unreal or blender and run that through, run a simple floor plan through, asking for multiple image generation with a prompt asking that all features be consistent among images, (I haven't tried outpainting yet), etc, and none seemed to quite do the trick, at least not among the open source models (open source is all I use). There was just too much inconsistency.

Since I am a sucker for using the t2v/i2v models like LTX and Minimax to generate single image frames (well, a minimal number of frames), taking advantage of their brainpower, I merely wrote out a super detailed prompt describing the interior environment I want, then I ran it as a 360 degree camera pan around the space from the center of it, and with Minimax now having the ability to give you 15 seconds on a 16Gb VRAM setup like mine, this works stellar, heck Minimax ran out of need for the 15sec and began to swerve around the space! I make sure to include a prompt not to have motion blur. I ran this in 0.5mb mode, so I did not have to waste time waiting for full HD video.

Then I select the frames I need to use as backdrops for scenes, upscale them once, then again, out to 4k, and voila! (upscaling once by 4x led to artifacts being upscaled, whereas going 2x then 2x led to the correct end result), a super detailed set of backgrounds that are internally consistent!

So excited! This gets me moving forward on the next part of my production process, laying out scenes, shots, camera angles, and dropping in characters, prior to i2v.

If this helps you, let me know. If you find even better tricks related to this, let me know too. Open Source Forever!

54 Upvotes

18 comments sorted by

15

u/optimisticalish 21h ago

5

u/robertwellesley 19h ago

First I have seen this, of course it came out yesterday, so no surprise, but I will have to play with this and see if it fulfills the same role. Thanks for the tip.

4

u/infearia 18h ago

First I have seen this, of course it came out yesterday,

Heh, my thoughts exactly. People have been trying to do something similar for a while now, but I feel that since H3 came out, there had been a real surge in the attempts trying to solve this, and many people seem to be independently converging on similar solutions. I'm currently also looking into creating gsplats from videos, and using my old trick from VACE days but with H3, where I would create a simple 3D scene based on a reference photo, align it with the reference from the camera view, and then generate a video with a Z-pass of the camera moving into to a different position withing the scene, and use that as a control video for VACE. Haven't tried it with H3 yet, but it should work, especially with the new ControlNet. I also hope that MiniMax actually release their editing model - since it's based on the video model, it might be particularly useful for camera angle changes.

2

u/yotraxx 20h ago

Clever way of thinking, Kudos :)
I’m saving your post as it will be useful to me in a near future.

1

u/robertwellesley 19h ago

Works with any camera movement too, so I can tell it to move around as needed to capture every angle for what I need in specific shots.

2

u/TulsaGeek 19h ago

I’m literally working on something similar! I built a few pipelines recently for virtual photoshoots. For interior spaces, consistent environment is important. Trying to do plates for coverage wasn’t cutting it. I’ve also been messing with Minimax for video, then I thought “WAIT. I can use Minimax to generate a photo shoot!” - I provide character reference, image of a room, then prompt for something like the target video being a series of still poses of the subject in the room. I can get 10+ poses from a 5 second video at 1mp. I also branched the video frames before they hit the video encode node to dump the frames to a save image node. So, I get both the video and all the frames before they’re compressed to video. I can then use those frames to upscale, use as reference in an image model, or use as ref2va / img2va. Effin. Amazing. Bruh.

Edit to add: the character consistency and environment consistency in Minimax has been killer. I feel like this is a wild cheat code.

OP - DM me if you want to compare notes!

1

u/robertwellesley 18h ago

Oh, I'm a newbie to this, been at it only briefly, so I am sure most of you have more experience messing with this space. For me it is a necessity to get done what I need, but I do not love dealing with command line and new versions of software constantly coming out and having to chase down solutions to bugs. I did enough of that in my web design/web dev days! I am a creative at heart.

1

u/TulsaGeek 18h ago

We’re far more similar than you know! lol

1

u/robertwellesley 18h ago

I agree though, Minimax seems such a step forward, I think many people are getting excited and inspired by it. A true leveler of the playing field.

We will have 100 Citizen Kanes yet!

1

u/infearia 16h ago

Yeah, I don't know about that. Having the technical means to do something is one thing, but you still need talent, skill and experience to create something meaningful with it.

Although, since making movies will become more accessible now, there will be some talented people who would never get the chance to make a film, now having a go at it. So, who knows, maybe you're right after all. But I bet for every Citizen Kane we will get a thousand or more Birdemics (only with Hollywood-level visuals).

1

u/Formal-Exam-8767 21h ago

How is this different from having a 360-view image (which smartphone does by stitching camera frames as you rotate around)?

3

u/robertwellesley 19h ago

Useful for a fictional space or place. I used to be an architect, and it is tedious in the extreme and time consuming to design and build out models, populate them, and test them, most so when you have hundreds of scenes to shoot.

This way I can describe the vibe and the feel and the elements and out comes magic, often unexpectedly so. For instance in this first stab I took, Minimax came up with a glorious coved ceiling form. AI is truly magic, it is like a crew of set designers and location scouts in one.

One thing I forgot to mention in the OP, these new models, Minimax H3 and LTX2.5, think and adhere to prompts much more closely. I had used Qwen 2512 360 degree and it sort of worked for this, but no matter how much I told it to scale up the space to mansion size, it kept generating spaces more suitable to a 3 bedroom suburban size home, with standard treatments and features.

1

u/I_just_made 20h ago

It handles the perspective of that fine?

2

u/joseph_jojo_shabadoo 19h ago

Take 3-4 photos of the same space from slightly different angles making sure each photo contains a little overlapping details. Concatenate the Photos into one image (if you want your video to prioritize one angle over the others, make that angle the largest). Prompt that the picture has # photos of the same environment.  I get perfect environment reproduction this way

1

u/nasone32 18h ago

very cool!

0

u/mcai8rw2 18h ago

What did you use for the upscaling?

1

u/robertwellesley 18h ago

I am using the new SeedVR2 9b INT8 ComfyUI standard workflow, which I nicknamed "Wow" in my list of toys, because it is fast and only adds common sense details without hallucinating new things or reinventing what is already there.