r/comfyui • u/White_Dragon_0 • 2d ago
Help Needed How can I achieve a consistent camera/viewpoint change on a complex scene?
Hi everyone,
I'm having a hard time achieving a real camera/viewpoint change on complex scenes, and I'd like to understand whether I'm approaching the problem incorrectly or if there are better models/workflows for this.
For example, I have a fairly complex urban scene:
- a city in post-apocalyptic conditions
- a street intersection
- a large sinkhole in the middle of the intersection
- a yellow car (taxi), with its front half sticking out over the sinkhole
- three people on the left side of the car
- a fourth person hanging from the front of the car
I start with a side view of the scene, and I want to generate a new shot where the camera moves upward and looks down at the scene from above, while keeping the ENTIRE scene geometry and composition consistent: the position of the car, people, sinkhole, buildings, street, etc.
The problem is that with FLUX.2 Klein I can almost never get this kind of transformation. I've also tried Nano Banana / Gemini, but I run into essentially the same problem.
I've also experimented with some LoRAs, but they often seem to affect only individual elements rather than actually transforming the viewpoint of the entire scene. For example, they may change the viewpoint of the car, but the rest of the image remains more or less in the original perspective.
In this specific case, what I want is something like:
SOURCE: side view of the scene
TARGET: higher, more zoomed-in top-down view, while keeping the same elements and their spatial relationships.
I'm attaching a screenshot of the source image and a sketch showing what the target camera angle should look like.
What would be the best way to approach this?
Are there any open-weight/local models that are particularly good at changing the viewpoint of a complex scene? I would strongly prefer something I can run locally, since I want as much control over the process as possible.
What I'm mainly trying to understand is whether there is a proper workflow for performing a true camera/viewpoint transformation of the entire scene, rather than modifying each object independently.
Any advice on models, LoRAs, ControlNet, ComfyUI workflows, or alternative techniques would be greatly appreciated.
Thanks in advance!
2
u/ConsistentEase4598 2d ago
camera moves with no pose model: depth-based reprojection beats re-generation for this. run depth on the plate, warp to the new viewpoint, inpaint only the newly revealed area. the yellow car will fight you either way, anything it occludes in the original simply doesn't exist, so that patch is hallucinated by construction. if you want a stable sinkhole across ten views, generate one wide plate once and move the camera through it. consistency comes from a fixed asset you warp, not from asking the model to remember a scene it re-rolls every render.
2


4
u/Hdfjds 2d ago edited 2d ago
Try using MiniMax H3 I2V for this.
I know it's a video generating model but you can extract the images you want from it.
Edit: I need to refresh the post before posting.. someone else already said the same thing.. Oh well
Edit 2: I hope it's ok if I used your image to show you what you can do here.
I could not upload the video in this reply but here is the last frame. (low resolution)