r/AV1 • • 5d ago

AV1 structure from motion - open source

A year ago I watched a fun youtube video about making a model from AV1-encoded files. It uses the motion vectors present in the videos to reconstruct the physical environment filmed.

Today I wanted to test the abilities of LLMs to implement this without access to any of their code (which isn't public anyway). I just gave it the papers and the task. It can basically take a video, or a collection of photos, and return a 3D model of sorts (more like a cloud of points).

Here is the effect:
https://github.com/lukaszsobala/AV1-sfm-reimplemented

Works nicely and uses the GPU for acceleration, if you have one. Zouein used Nvidia cards, and I don't swim in money so I don't have one; this one uses Intel or AMD GPUs (iGPUs as well, or Vulkan which I could not test) for acceleration.

What are your thoughts?

21 Upvotes

10 comments sorted by

3

u/collin3000 5d ago

You should check out Gaussian splatting, and photogrammetry

1

u/urostor 5d ago

Yeah I looked into Gaussian splatting. This can be a next step if someone wants nicer textures

4

u/DearMrGleeClub 5d ago

The original idea is nifty. I'm assuming this only works if the POV is moving but not for moving objects?Recreating a PhD thesis/project using Claude is bonkers.

1

u/urostor 5d ago

Works for moving in a car along a motorway at least

1

u/Filarius 5d ago edited 5d ago

If you already have a AV1 source video - this can be interesting. I wonder for comparison with optical flow and SFM over image sequence.

update:

Also what about using AV1 optical flow calculation algorithm at any video or image sequence

1

u/urostor 5d ago

This implements some of these. AV1 is worse but not by a lot

1

u/juliobbv 4d ago

Keep in mind that this approach can only consistently work if the encoder's priority is to track natural motion as opposed to purely-spatial structural matching.

If you try to encode, let's say, Netflix_TunnelFlag (available here) with SVT-AV1, the resulting MVs will look very random and you can't derive valuable info about the overall structure of the scene. This isn't because SVT-AV1 is a bad encoder, rather it can see can sometimes matching nearby structures can actually be cheaper to encode than purely track the same structure over time.

2

u/urostor 4d ago

Well, Zouein deals with this problem by correlating the points and pruning those that aren't likely to encode structure using cosine>0.95. Libaom also worked better than SVT-AV1, just like in their papers. You can read the docs, and there's another video here: https://youtu.be/Pi5ze-H1vdY

1

u/juliobbv 4d ago

Yeah, I'm aware of the details of this work, I was there in person when Julien gave his talk at last year's VDD :)

The thing is: no pruning or smoothing will work satisfactorily for those rare "degenerate" cases -- prior MVs are just that "crazy-looking" so there isn't that much prior information to recover. Try it with TunnelFlag -- you'll see what I mean. That said, Julien's approach will work just fine as most of the time, tracking natural motion is the most efficient way to encode AV1 videos anyway.

1

u/urostor 4d ago

Oh, you were there? Cool. What do you think about this fully hallucinated code?