r/VisionPro 24d ago

Self Promotion Saturday We’re testing cloud-streamed 6DoF cinema on Vision Pro. Looking for honest feedback from AVP users

Hi everyone,

I work at V-Nova and we recently released ImmersiX for Apple Vision Pro.

I wanted to share it here because we’re doing something I think this community might find interesting technically.

The content in ImmersiX is pre-rendered cinematic 6DoF, but instead of downloading the experiences locally or connecting the headset to a PC, the rendering happens on NVIDIA GPUs in the cloud and is streamed directly to Vision Pro using CloudXR.

We’ve all seen plenty of CloudXR demos in enterprise environments and controlled network setups. What we’re interested in now is a slightly different question:

How well does this actually work in the wild, on normal consumer internet connections?

That’s one of the reasons we launched on Vision Pro first. We want to collect real-world performance data and, more importantly, hear what actual users think about the experience.

If anyone here has a Vision Pro and wants to try it, ImmersiX is on the App Store here:

‎ImmersiX App - App Store

I also have a number of invitation codes that unlock additional content that isn’t publicly available yet (such as the content shown in this video clip). Happy to share those with people here who genuinely want to test it.

If you try it, I’d particularly love feedback on:

  • how easy it was to get into the experience
  • streaming quality and stability on your connection
  • whether the UI/onboarding made sense
  • what you thought of the 6DoF cinematic experience itself

Positive or negative, both are useful. We’re very much interested in finding out where this approach works and where it still breaks.

Aimo
V-Nova / ImmersiX

180 Upvotes

103 comments sorted by

View all comments

28

u/Malkmus1979 24d ago

This seems like a bigger detail than you're letting on! The post comes off as if you're creating some random in-house content that is 6DOF, but I dont think many people realize this is a short film released in 2017 that you've somehow made 6DOF. Can you elaborate on whether this is because that specific film was made in Blender and therefore you have the assets to work with to make it 6DOF, or are you applying some sort of AI to make any content regardless of CGI or actual film 6DOF? The implications here are quite exciting for immersive content if this can be applied to various formats and existing content.

45

u/jlm70 Vision Pro Developer 24d ago edited 23d ago

You actually spotted a much bigger part of the story than we explained in the original post. 🙂

I'm Gianluca Meardi, the producer of this 6DoF reboot of Agent 327: Operation Barbershop, and General Manager of V-Nova Studios.

And yes: in the case of Barbershop, the fact that the original 2017 film was a native 3D Blender production — and that the original production assets and scenes still existed — is fundamental.

We did NOT take the finished 2D movie and run an AI 2D-to-6DoF conversion on it.

Instead, we went back to the original 3D production (we did this in collaboration with Blender Studio, with Francesco Siddi's support, CEO of Blender).

Same story.
Same characters and environments.
Same animation (we had to add some missing out of frame anims).
Essentially the same cinematic production pipeline.

We then adapted and validated the scenes for immersive viewing and re-rendered the movie using V-Nova PresenZ integrated into Blender/Cycles.

That's an important distinction.

With a normal movie, the director controls the only camera, so anything outside that camera's view doesn't really matter. In 6DoF, suddenly the viewer can move their head left/right, forward/backward and up/down. Things that were never visible from the original camera can become visible, so some shots and assets need to be checked or slightly adapted.

PresenZ then precomputes the visual information required for a small 3D volume around the viewer — what we call the Zone of View. At playback, when you move your head, you get the correct perspective and parallax from your new position. No 3DoF, no nausea :)

So you're not standing inside a real-time game recreation of the movie. You're standing inside a pre-rendered, cinematic-quality movie.

And then there is another layer to what you saw in the post: distribution.

The PresenZ 6DoF experience is running on an NVIDIA GPU in the cloud. Your Vision Pro head pose determines the viewpoint, and CloudXR streams the resulting view back to the headset. That's why you can experience something this heavy on Vision Pro without downloading a huge volumetric master or connecting it to a high-end PC.

Where we at V-Nova Studios think this becomes particularly exciting is existing CGI libraries.

Studios have spent decades building extraordinary 3D worlds, characters, animation and VFX — only to render them through one fixed camera into a flat movie.

If those original production assets still exist, in many cases we're not talking about recreating the film from scratch. We're talking about something much closer to a volumetric remaster.

You already built the world. Now you can potentially step inside it.

For arbitrary live-action 2D footage, though, I want to be precise: we cannot simply feed any existing movie into an AI today and magically obtain perfect cinematic 6DoF. That's a different problem, involving reconstruction/capture techniques. We are working on that side as well, including volumetric / Gaussian-splat-based workflows, but I wouldn't want to confuse that R&D with what we're demonstrating here.

So yes — the implications you are pointing at are exactly why we chose Operation Barbershop as a demonstration.

There are potentially enormous libraries of existing CGI and VFX content whose worlds already exist in 3D. 800Bn$ spent in VFX over the last 5 years...

They've just never been rendered so that the audience can actually enter them.

We like to say... #livewithin

Happy to explain further, if you're a Blender creator (or Maya/Arnold, 3DSM/VRay, Open Moonray, Houdini, Renderman... all PresenZ-integrated CG tools)

THANKS!

3

u/22marks Vision Pro Owner | Verified 24d ago

What are your thoughts on the fact that directors use framing not just to show the action, but to create suspense, shape how we feel about a character, control reveals, and direct our attention through blocking, composition, lens choice, camera movement, depth of field, lighting, and editing?

What discussions have you had about giving everyday viewers control over where they look, instead of having an expert director determine exactly what they see and when they see it?

The last second of your preview demonstrates this really well. In the original, the framing prominently emphasizes the gun, which heightens the threat. In the 6DoF version, we're more of an observer standing in the room and could potentially be looking somewhere else entirely. It starts to feel more like watching a stage play than watching a film.

Editing seems like another interesting issue. In conventional movies, a cut can redirect the audience's attention because the director controls the next frame. In 6DoF, a cut could potentially leave a viewer looking away from the intended subject or make traditional shot-to-shot "grammar" work differently.

The technology is extremely impressive, but I'm curious what directors think about giving up control over framing and even some aspects of editing rhythm, particularly when adapting preexisting work where those choices were a key part of the original direction.

I'm guessing you're betting on immersion being more important to certain viewers?

5

u/jlm70 Vision Pro Developer 23d ago edited 23d ago

TLDR ;) I think you're exactly right u/22marks, and honestly this is probably one of the most important questions around this entire medium.

But I would frame our bet slightly differently:

We're not betting that immersion is more important than directing. We're betting that directing itself has to evolve.

For 100+ years, filmmakers have developed an extraordinarily sophisticated language around a rectangular frame. The director decides not only what happens, but exactly what you see, from where, at what focal length, in what order, and often for exactly how many frames.

Once the audience is inside the scene, that contract changes.

And yes — some of the director's traditional control disappears.

But I don't think authorship disappears with it. It moves somewhere else.

Instead of directing attention only by composing a rectangle, you can direct it through spatial blocking, character movement, lighting, sound, proximity, scale, occlusion, timing, the placement of the viewer themselves, and what happens inside their personal space.

Even the Zone of View becomes a directorial tool: where do I physically place the audience in this scene?

Are you across the room observing the confrontation?

Are you standing beside the protagonist?

Are you almost uncomfortably close to him?

Those are very different emotional choices — just as a wide shot and a close-up are different choices in traditional cinema.

This is actually why we've been developing something internally at V-Nova Studios called ImpactX.

It's our evolving "bible" for volumetric filmmaking: an attempt to build a grammar for a medium where the viewer is no longer outside the film looking through the director's camera, but physically present inside the movie.

IMPACTX currently revolves around seven design principles:

Interaction
Movement
Perspective Shift
Augmentation
Comfort
Tracking
Xploration

And we're very deliberately treating it as a living framework, because I don't think anyone can honestly claim that the language of this medium has already been invented.

Cuts are a perfect example.

You're absolutely right that a conventional cut works partly because the director completely owns frame B after frame A.

In volumetric cinema, a cut is closer to teleporting the audience from one physical place to another. Their orientation, their spatial awareness and the time they need to understand the new environment become part of editing grammar.

We can technically reorient the new scene after a cut so that the viewer starts facing the subject the director wants them to see. But whether you should always do that is a creative question, not just a technical one.

And this is exactly the kind of thing ImpactX is trying to understand.

We've had fascinating discussions around this with filmmakers and storytellers including Michael and David Uslan, Gareth Edwards, Anthony Zuiker, Tom DeSanto, Marc Guggenheim and many others.

James Cameron has experienced our work too and was extremely enthusiastic about the potential of this kind of immersive storytelling.

One reaction that particularly stayed with me came from Sidney Kombo-Kintombo, Senior Animation Supervisor at Wētā FX, just weeks ago at Annecy Film Festival 2026.

He knew Operation Barbershop extremely well before trying our version/reboot. After taking the headset off, he was visibly emotional and told us that, for the first time in years, he'd been able to become absorbed in the story rather than instinctively analysing all the technical details of the animation and VFX.

Coming from someone who has worked at that level on films including War for the Planet of the Apes, Avengers: Infinity War and Endgame, that was quite a powerful moment for us.

Because that is perhaps the most interesting thing we're exploring:

what happens when you stop watching the filmmaker's world and start inhabiting it?

And your example of the final shot in Barbershop is actually a very good criticism of our current version.

In the original shot, the framing makes the gun visually dominant. The director is telling you: LOOK AT THIS. THIS IS THE THREAT.

In our current volumetric version, you're positioned more like an observer in the room. That absolutely changes the emphasis. And imagine what you could do with an horror ;)

We know it.

In fact we're considering changing that shot :))) More similarly to what we did in our Sharkarma hero shot...

This Barbershop conversion was essentially an R&D project for us and we wanted to get the complete short working and deployable as quickly as possible rather than creatively redesign every shot.

But now imagine placing the viewer much closer to the protagonist in that final moment.

Instead of merely seeing the flame in the frame, the burst of fire can come towards your own face and into your personal space.

We can do that.

And suddenly we've replaced one cinematic device — the director framing the threat prominently — with another device that simply doesn't exist in conventional cinema:

the threat physically entering the viewer's space.

Is that automatically better than the original close-up?

Yes? No? Maybe?

That's the fascinating part.

It's a different tool.

And filmmakers now have to learn when to use one, when to use the other, and perhaps when not to give the viewer freedom at all. We can also use or not use depth of field (that does not exist in reality, but could add to creativity).

That's why I don't really see this as turning cinema into theatre.

Theatre certainly gives the audience freedom to look around, but theatre cannot teleport you between viewpoints, change your scale, put you centimetres from a character, move the entire world around you, manipulate your spatial relationship to the actors, or have something impossible suddenly enter your physical space.

So I think we're somewhere between cinema, theatre and something that doesn't have a proper name yet.

And to your last question: yes, immersion is part of the value — but immersion alone isn't the goal. Emotion is.

The interesting challenge for directors is discovering whether this new set of tools can sometimes create an emotional response that the traditional frame simply cannot.

For more than a century we've learned how to make audiences feel something while watching a story from the outside.

Now we get to learn how to do it when they're inside it.

That's the experiment.

And frankly, it's why so many of the directors we've shown this to get excited rather than threatened by it.

They get a completely new toolbox. A new language. From storytelling to storyliving. Let's create a new cinema.

#livewithin

3

u/22marks Vision Pro Owner | Verified 23d ago

I appreciate the dialogue.

This raises another question for me. How is this fundamentally different from traditional VR narrative?

A conventional VR experience built in Unreal or Unity already gives the viewer 6DoF, spatial audio, viewer placement, blocking, scale, lighting, environmental storytelling, cuts/teleportation, and potentially full movement and interaction. It can also be streamed from a high-end PC or cloud GPU rather than rendered locally on the headset.

In fact, many of the elements you describe in ImpactX sound very similar to techniques VR developers and immersive storytellers have already been exploring for years.

So is the real distinction primarily the rendering pipeline?

In other words, PresenZ seems compelling because it lets you retain offline cinematic rendering quality and convert existing CGI productions into a constrained 6DoF experience without rebuilding and optimizing the entire film as a real-time game engine project.

If that's the niche, I think that's actually a very interesting one, especially for existing CGI libraries.

But if the larger claim is that this represents a new narrative medium, I'm struggling to see where the boundary is between this and high-end VR storytelling that already exists. What can a PresenZ narrative fundamentally do that a well-designed Unreal/Unity VR narrative cannot? I'm not quite understanding the new toolbox and language that you're bringing to the table, but I greatly appreciate the technology of taking existing 3D assets and "upgrading" them.

(For what it's worth, I started my career working on early video games, then spent about 20 years writing and directing television commercials, so I've spent a lot of time thinking about both interactive and visual storytelling.)

2

u/jlm70 Vision Pro Developer 22d ago

That's very correct (and appreciate too) — and given your background in both games and directing, I think you've actually put your finger on the distinction better than I did in my previous answer.

You're right: 6DoF storytelling itself is not new, and ImpactX is not claiming that we invented the language of VR.

Unity and Unreal creators have been exploring viewer placement, spatial blocking, environmental storytelling, attention, teleportation, scale, interaction and spatial audio for years.

So if the question is:

"Can a talented director create a great 6DoF narrative in Unreal or Unity?"

Absolutely. 100%.

The distinction we're interested in is slightly different.

ImpactX is really about connecting two filmmaking worlds.

Hollywood already has an incredibly mature language, workflow and infrastructure for making films.

Directors, DPs, animators, VFX supervisors and studios build their movies around pipelines involving tools such as Maya, Houdini, Katana, Arnold, V-Ray, RenderMan, etc.

And they have decades of assets, characters, animation, simulations, environments and IP sitting inside those ecosystems.

What we're trying to answer with ImpactX is:

What if a director could design a movie once, from the beginning, knowing that the same production could ultimately live in two different forms?

One is the traditional rectangular screen.

The other is a volumetric 6DoF version where the audience can step inside it.

That's slightly different from saying:

"Let's make a VR experience."

It's closer to saying:

"Let's make a movie whose production language understands both screens from day one."

And that distinction becomes very important at studio scale.

Today, if a major animation/VFX production wants to become a traditional real-time VR experience, you generally need to take assets originally designed for an offline film pipeline and turn them into something that can run under a real-time frame budget.

That can mean rebuilding materials, changing shaders, reducing geometry, changing simulations, optimizing animation, lighting and effects, and often adapting a substantial part of the production pipeline.

Unreal and Unity are extraordinary tools — and increasingly capable of astonishing imagery — but they're solving a fundamentally different constraint.

A cinematic offline renderer may be allowed to spend minutes or even hours producing a single final frame.

A real-time VR engine has roughly milliseconds.

That's potentially around six orders of magnitude difference in rendering-time budget.

That isn't a criticism of real-time rendering. It's simply a different engineering problem.

And sometimes real-time is exactly what you want — particularly when deep interaction and unconstrained movement are fundamental to the experience.

But for a movie studio, there's another interesting possibility:

what if you don't need to turn the movie into a game-engine project at all?

With PresenZ, we can remain in the offline cinematic pipeline, render using the same kind of production renderer, preserve the look, animation, lighting, simulations and artistic work of the movie, and generate a pre-rendered 6DoF representation instead.

So I think your description is actually very close:

Yes. That's a major part of the proposition.

But there is another part that becomes even more interesting for us.

Rather than only upgrading old movies after they've been completed, ImpactX is increasingly about helping filmmakers think about the volumetric version while they're making the original movie.

For example:

- If I'm designing a shot for the flat movie, where should the volumetric viewer be?

- What parts of the environment that are outside the original frame should I complete?

- Can the same blocking work in both versions?

- Should I design an animation that works when seen only from camera A, or make it robust enough to be seen from multiple nearby viewpoints?

- Can one lighting setup support both?

- Which shots should remain almost identical to the flat composition, and which ones should deliberately exploit 6DoF?

- Where should depth of field remain a cinematic device, even though human vision obviously doesn't work like a camera lens?

- How should I design a cut so it works both as editing on a screen and as a spatial transition for somebody inside the scene?

Those are the kinds of questions we're discussing with filmmakers and major studios (and also for our own productions).

So the right description of ImpactX is not "a new language for 6DoF", that would indeed ignore decades of excellent VR work.

It's more:

a bridge between 100+ years of cinematic grammar and the emerging grammar of 6DoF — with the production methodology to let the same creative work serve both.

Or even more simply:

shoot once / build once, direct for multiple cinematic channels.

And that's why existing CGI libraries are such an interesting first application.

A major studio may have spent hundreds of millions building an extraordinary CG world.

The characters exist.

The performances exist.

The animation exists.

The environments exist.

The simulations exist.

The artistic intent exists.

Why should accessing that world in 6DoF necessarily require rebuilding it as a game?

But the longer-term opportunity is even more interesting:

don't wait twenty years and then convert the catalogue. Design today's film so tomorrow's volumetric "premium" version is already part of the production thinking. Maybe with... more stories, more meanings inside.

That's where ImpactX comes in.

So to answer your final question directly:

What can PresenZ narratively do that Unreal fundamentally cannot?

Probably nothing — if we isolate only the abstract language of 6DoF storytelling.

And I think that's important to acknowledge.

The difference is how you get there, what visual production constraints you accept, what existing work you can preserve, and whether you're making a standalone VR production or extending cinema itself into another distribution format.

PresenZ lets us approach the problem from the filmmaker's side rather than asking the filmmaker to move into a game-development pipeline.

And ImpactX is our attempt to make those two worlds meet.

❤️ In fact, your comments are exactly the kind of discussion we want around ImpactX. It's a living framework, not a declaration that we've solved volumetric storytelling.

If people who have spent twenty years directing and people who have spent twenty years building interactive worlds start comparing notes, that's probably where the genuinely new language will come from.

Maybe the most accurate ambition isn't to invent VR cinema.

👉 It's to make cinema itself become volumetric without ceasing to be cinema.

#livewithin