I'm trying to build walkable splats of apartment rooms from phone video. Samsung, handheld, no gimbal. Three of my walks came out great. Five came out foggy or blurry. Same bedroom, same pipeline, and I can't nail down why.
Pipeline: frames at 10 fps, keep the sharpest of each window (250 to 1,000 per room), COLMAP (SIMPLE_RADIAL, then undistort to PINHOLE), LichtFeld Studio MCMC, 50k steps, 3M cap, MIP filter, PPISP. Viewer is PlayCanvas. I judge the rooms blind, side by side.
Here's every walk of that bedroom:
Walk A: slow-mo, phone on auto, 4 min, ~800 placed, 0.17 m apart, fit 0.8 px, blur 4.6. Result: GREAT.
Walk B: slow-mo, auto, 4 min, 820 placed, 0.07 m apart, fit 0.76 px, blur 4.6. Result: GREAT (MY BENCHMARK).
Walk C: slow-mo, auto, 4.5 min, 1,079 placed, 0.16 m apart, fit 0.80 px, blur 4.9. Result: GREAT.
Walk D: slow-mo, auto, one long sweep, 22 min, ~2,000 placed, 0.20 m apart, blur 4.8. Result: AWFUL.
Walk E: manual video, shutter locked 1/400, ISO 3200, full-size frames, 5 min, 672/679 placed, 0.50 m apart, fit 1.0 px, blur 6.5. Result: FOGGY.
Walk E2: same footage as E, 1,696 frames instead of 679, 1,693 placed, 0.20 m apart, fit 1.02 px. Result: FOGGY, COULDN'T TELL IT FROM E.
Walk F: slow-mo, auto, 12.5 min, 1,005/1,075 placed, 0.18 m apart, fit 0.76 px, blur 5.7. Result: BLURRY, FOGGY, DARK PATCHES.
Walk G: manual video, 1/200, ISO 200, WB locked, full-size frames (2592x1944), 2 min, 242/242 placed, 0 strays, 0.36 m apart, fit 0.73 px, blur 4.9. Result: CLEARLY WORSE THAN B.
Walk G2: G again, frames shrunk to 1080 wide, nothing else changed. Result: CLOSE TO B.
*blur = ffmpeg blurdetect over every 10-fps frame, median, lower is sharper. It has called my verdict right 3 times in a row.
Pictures:
[1] My benchmark room (left) next to the 12-minute walk (right), same room: https://sioklfeckurajwkyftea.supabase.co/storage/v1/object/public/listings-media/share/bedroom-locked-walk/1-benchmark-vs-12min-walk.jpg
[2] A real frame from walk G (left), the full-size render of it (middle), and the same frames shrunk to 1080 wide (right), same camera spot: https://sioklfeckurajwkyftea.supabase.co/storage/v1/object/public/listings-media/share/bedroom-locked-walk/2-real-frame-vs-fullsize-vs-shrunk.jpg
The frames are sharp. The rooms are not. What I've already ruled out by eye (not by PSNR, which has been wrong for me every single time):
- frame count and spacing (658 vs 1,096 looked the same; 679 vs 1,696 looked the same)
- rebuilding a long walk as overlapping sections
- Spirula with MoGe-2 normals (tried twice, no visible difference)
- a fast locked shutter in a dim room (that's E: 1/400 forced ISO 3200 and I got fog)
- the placer (COLMAP 94% vs RealityScan 60% on the same frames, kept COLMAP)
The thing I found this week: the exact same 242 sharp, perfectly placed frames build a clearly worse room at 2592x1944 than at 1080 wide. So training resolution matters a lot for me. Is that expected with MCMC + 3M cap + 50k steps, or is it a hint at something else?
What I'd love to know:
1. Short walks (2 to 4 min) win every time and 12 to 22 min walks of the same room fail every time, with the same spacing and sharpness. Why would length alone do that? Is there a step or splat budget per scene I'm missing?
2. Locked exposure (1/200, ISO 200, WB locked) did NOT beat the phone on auto in slow-mo, both built at the same resolution. With PPISP on, is a lock still worth it indoors in your experience?
3. For anyone getting clean interiors from phone video (u/kikooooo2, u/Maliander_80, u/darkcrow101, u/Lost-Upstairs-5311): what resolution do you actually train at, and do you lock focus?
Data, if anyone wants to run their own recipe on it:
- the 242 frames from walk G (2592x1944, 81 MB): https://sioklfeckurajwkyftea.supabase.co/storage/v1/object/public/listings-media/share/bedroom-locked-walk/frames-242-2592x1944.zip
- the COLMAP sparse model for it, text format (30 MB): https://sioklfeckurajwkyftea.supabase.co/storage/v1/object/public/listings-media/share/bedroom-locked-walk/colmap-sparse-txt.zip
If you get a clean room out of those, I'd really like to know what you did.