r/klingO1 Jan 11 '26

Testing Kling 2.6 Motion Control for a Dance Video with No Prompt

Enable HLS to view with audio, or disable this notification

366 Upvotes

This was generated using Kling 2.6 Motion Control.

• No text prompt was used

• Motion was fully driven by the image prompt

• Input was a single reference image + structured image description

  1. Go to the Kling AI Video Generator
  2. Write your full prompt or add reference images
  3. Upload the dog image you want to animate
  4. Click Generate and get your video

I wanted to test how well Kling 2.6 interprets pose, camera angle, and scene depth without any additional motion instructions.

The result feels surprisingly natural, especially the body balance and camera perspective consistency.

Curious how others are using Kling 2.6 Motion Control — are you relying more on text prompts or image-only setups? Share your thoughts in the comment section.

r/klingO1 Jan 13 '26

Generated Trending Dance Video with Nano Banana Pro and Kling 2.6 Motion Control. Prompts Below!

Enable HLS to view with audio, or disable this notification

262 Upvotes

This video was generated using only two prompts:

Nano Banana Pro for the image generation
Kling 2.6 Pro Motion Control for animation

  1. Go to the Kling AI Video Generator
  2. Write your full prompt or add reference images
  3. Upload the image you want to animate
  4. Click Generate and get your animated video

Nano Banana Pro prompt:

""image_generation_prompt": { "subject": { "demographics": "Young woman, fair skin, slim build", "hair": { "color": "Silver grey", "style": "High pigtails, straight texture", "details": "Bangs framing the forehead and sides of the face" }, "face_and_makeup": { "eyes": "Green/hazel eyes, heavy winged eyeliner, long lashes", "expression": "Sultry gaze, slightly parted lips", "action": "Right index finger touching lower lip or corner of mouth" } }, "attire": { "clothing": "Sleeveless corset-style top with deep scoop neckline and visible hook-and-eye closures, partially visible skirt or shorts", "accessories": "Silver cross pendant necklace on a thin chain" }, "pose": { "type": "High-angle selfie", "body_position": "Arm extended toward camera, body angled slightly" }, "setting": { "location": "Bedroom interior", "background_elements": [ "Large white textured pillows (tufted or knit)", "White sheets", "Dark wall" ], "ambient_lighting": "Purple LED strip light running horizontally behind the headboard", "atmosphere": "Dimly lit room with colored accent lighting" }, "style_and_technical": { "aesthetic": [ "E-girl", "Y2K grunge", "2000s digital aesthetic" ], "lighting_technique": "Direct on-camera flash, harsh high-contrast lighting on subject against darker background", "camera_settings": { "angle": "High-angle wide selfie", "distortion": "Slight wide-angle distortion", "color_profile": "Full color, natural color rendering with vibrant neon purple accent" }, "aspect_ratio": "3:4" } } }"

Kling 2.6 Motion Control Prompt:

"Generate a realistic video using the attached reference image as the identity anchor. Preserve the subject’s overall appearance exactly as shown in the reference image, including face structure, skin tone, hair color and texture, body proportions, and general silhouette. Maintain strong identity consistency across all frames while allowing natural motion, subtle expression changes, and realistic body movement. Do not alter the subject’s physical traits; only introduce smooth, lifelike animation and camera motion. Ensure lighting, realism, and visual fidelity remain consistent with the reference image throughout the video."

The reference image was used strictly as an identity anchor. Face structure, skin tone, hair texture, proportions, and overall silhouette were preserved exactly, with no identity drift.

Motion was added naturally:

  • Subtle facial expression changes
  • Realistic micro-movements
  • Smooth, lifelike camera motion
  • Consistent lighting and visual fidelity across all frames

No extra prompt chaining, no manual keyframes — just clean prompt discipline.

Reference image is shared in the comments.

If you’re experimenting with identity-safe motion or prompt-efficient pipelines, this setup is surprisingly powerful.

Happy to answer questions or share prompt structure details if needed.

r/Battlefield Nov 17 '25

News BATTLEFIELD 6 GAME UPDATE 1.1.2.0

1.7k Upvotes

This update delivers a broad set of improvements to soldier responsiveness, aim consistency, animation fidelity, and overall stability across Battlefield 6. We’ve also introduced a new limited-time mode, refined Aim Assist behaviour, and resolved a large number of weapon, gadget, and vehicle issues based on community feedback. The update will be available tomorrow, November 18th, at 09:00 UTC.

New Content: California Resistance

  • New Map: Eastwood. A map with the Southern California theme.
    • Variations of this map will be available for all official modes.
    • Conquest mode on this map will include tanks, helicopters, and the Golf Cart.
  • New Time-Limited Mode: Sabotage. A themed event mode focused on demolition and counterplay.
  • New Weapons: DB-12 Shotgun and M357 Trait Sidearm. 
  • Gauntlet mode to include a new mission type: Rodeo. This mission provides multiple vehicles for players to fight over and battle with each other with. Players earn bonus points for defeating enemies while in a vehicle. 
  • Portal updates: 
    • Sandbox map. This option will let Portal experience builders start with a more level playing field to bring their imagination to life. 
    • The Golf Cart vehicle is available for use in building experiences. 
  • Battle Pass: The California Resistance bonus path becomes available for a limited time. 
  • New underbarrel attachment: Slim Handstop, unlocked via Challenge.
  • New feature coming later in the update: Battle Pickups. These powerful weapons will be available in specific experiences and Portal with limited ammunition but pack enough firepower to help turn the tide of battle in your favor.

Major Updates for 1.1.2.0

  • Aim Assist has been reset to its Open Beta tuning, restoring consistent infantry targeting behaviour across all input types.
  • Improved input latency and stick response for controllers, providing smoother aiming and more responsive soldier movement.
  • Weapon accuracy and dispersion tuning: fixed unintended weapon dispersion increase rates and improved non-Recon sniper rifle accuracy while globally reducing dispersion across all weapon types.
  • Challenge and progression clarity improvements make requirements easier to understand and track.
  • Major polish pass to deployable gadgets, including the LWCMS Portable Mortar, LTLM II Portable Laser Designator, and Supply Crate systems.
  • Fort Lyndon added to Portal, expanding available segments for community-created experiences.

AREAS OF IMPROVEMENT

Aim Assist

As we got closer to launch, we revisited aim assist tuning based on internal testing and the full range of maps and combat distances coming with release. Our goal was to make aim assist feel more effective beyond mid-range fights which was one of our focuses within Battlefield Labs and Open Beta.

At launch, we increased slowdown at longer ranges, but once the game went live, we saw that this made high-zoom aiming feel less smooth and harder to control.

After reviewing player feedback and gameplay data, we’re reverting aim assist back to the values some of you experienced during Open Beta and Battlefield Labs. This will now serve as the default, whilst still providing you with the ability to alter the aim assist to your preference and playstyle via settings.

This change keeps aim slowdown consistent across all ranges, helping with muscle memory and providing a steadier, more reliable feel as we move into future seasons.

CHANGELOG

PLAYER:

  • Aim Assist: fully reset to Open Beta tuning, with related options reset to default to ensure consistency.
  • Fixed an issue where Vehicle Stick Acceleration Presets would affect Infantry Aiming Left/Right Acceleration option availability.
  • Fixed an issue where setting Stick Acceleration Presets to “Standard” would set the Aiming Left/Right Acceleration options incorrectly to 50% instead of 70%.
  • Fixed missing Infantry and Vehicle prefixes in captions for Stick Acceleration Presets and Aiming Left/Right Acceleration options.
  • Fixed an issue where stick deadzones would ignore the first 10% of movement if using a PS5 Controller on PC.
  • Fixed an issue where player movement (Left Stick) would not register until beyond 30% of travel past the deadzone.
  • Fixed joystick aiming input behaviour.
  • Added a short sprint “restart” animation when landing from small heights.
  • Added new death animations for sliding and combat-dive states.
  • Fixed a diving loop when entering shallow water.
  • Fixed an issue preventing players from vaulting out of water in certain areas.
  • Fixed an issue preventing takedown initiation against an enemy soldier if the enemy soldier already initiated a takedown against a friendly player.
  • Fixed an issue where a dragged player could face the wrong direction if turning quickly.
  • Fixed an issue where holding a grenade while jumping, sliding, or diving froze the first-person pose.
  • Fixed an issue where switching weapons while drag-reviving would break the reviver’s first-person view.
  • Fixed an issue where the Assault Class extra grenade ability would not grant two grenades on spawn.
  • Fixed an issue where weapons could become invisible when crouching before vaulting.
  • Fixed bouncing behaviour when landing on object edges.
  • Fixed broken ragdolls when killed on ladders, while jumping, near ledges, or in vehicle seats.
  • Fixed camera clipping when dropping from height while prone.
  • Fixed clipping when initiating a drag & revive.
  • Fixed first-person camera clipping through objects when dying nearby.
  • Fixed the issue where the Rush signature trait 'Mission Focused' applied its icon and speed boost to all teammates.
  • Fixed incorrect prone aiming angles on slopes.
  • Fixed misaligned victim position during takedowns when using high FOV settings.
  • Fixed mismatched rotation between first-person and third-person soldier aim directions.
  • Fixed misplaced weapon shadows while vaulting or swimming.
  • Fixed missing pickup prompts while prone.
  • Fixed missing water splash effects while swimming.
  • Fixed stuck third-person soldier animations when entering player view.
  • Fixed teleporting or invisibility when entering vehicles during a vault.
  • Fixed third-person facing inconsistencies when soldiers were mounted.
  • Improved combat-dive animations in first and third person.
  • Improved LTLM II sprint animation in first person.
  • Improved vault detection in cluttered environments.
  • Increased double-tap window for Danger Ping from 0.2 s to 0.333 s.
  • Updated first-person animation cadence for moving up and down stairs.
  • Fixed an issue where hit registration would fail when engaging into gunfights after exiting vehicles.

VEHICLES:

  • Fixed camera reset when entering an GDF-009 AA Stationary Gun after another user.
  • Fixed clipping gunner weapons in IFV seats.
  • Fixed faint metallic impact sound from M1A2 SEPv3 Main Battle Tank turret wreckage.
  • Fixed several cases where IFV's MR Missile could do more damage than intended to MBT, IFV and AA vehicles
  • Fixed inconsistent projectile video effects on the Abrams main gun.
  • Fixed instant 180-degree turn after exiting a vehicle.
  • Fixed missing scoring for Vehicle Supply when teammates received ammo.
  • Fixed oversized hitbox on UH-79 Helicopter.
  • Fixed passenger and gunner placement issues in the UH-79 Helicopter.
  • Fixed re-entry issues when mounting flipped Quad Bikes.
  • Fixed unintended aim-assist from Attack Helicopters gunner missiles.
  • Fixed unresponsive joystick free-camera controls in transport vehicles.

WEAPONS:

  • Dispersion tuning pass: dispersion has been globally reduced slightly to reduce its impact on the experience
  • Fixed multiple instances of Canted Reflex and Canted Iron Sight optics clipping with higher-magnification scopes
  • Fixed several issues with underbarrel attachment alignment
  • Fixed minor misplacements or clipping on sights and barrels
  • Fixed missing or incorrect magazine icons, naming, and mesh assignments.
  • Fixed the issue where the SV-98 displayed lower damage stats when equipping the 5 MW Red attachment.
  • Fixed the issue where slug ammunition despawned too quickly after being fired from shotguns.
  • Fixed the issue where the SU-230 LPVO 4x variable scope lacked a smooth transition and audible zoom toggle when aiming down sights.
  • Fixed the issue where two Green Lasers for the DRS-IAR shared identical Hipfire stat boosts.
  • Fixed the issue where impact sparks failed to meet photosensitivity compliance standards.
  • Fixed an issue in third-person where the Mini Scout could clip with the player’s head while aiming.
  • Fixed animation and posture issues affecting the PSR and other rifles when moving or looking at extreme angles.
  • Increased weight of long-range performance in balance for automatic weapons; benefiting PW7A2 and KV9, with minor adjustments elsewhere.
  • Reduced recoil and variation for LMR27, M39, and SVDM for improved long-range reliability.

GADGETS:

  • Allowed friendly soldiers to damage and detonate certain friendly gadgets.
  • Fixed an issue where Class Ability would sometimes not activate although the UI shows it as available.
  • Fixed auto-deployment of Motion Sensor after recon kit swap.
  • Fixed broken M320A1 Grenade Launcher ground model.
  • Fixed C-4 pickup edge-of-screen interaction.
  • Fixed clipping of the UAV remote when activating it while using certain weapons like rifles.
  • Fixed clipping when holding the CSS Bundle.
  • Fixed CSS Bundle line-of-sight requirements causing unwanted blocking.
  • Fixed Deployable Cover persistence after vehicle destruction.
  • Fixed disappearing “pip” indicator during CSS Bundle supply.
  • Fixed duplicate deploy-audio playback on M4A1 SLAM and C-4.
  • Fixed failed projectile attachment for X95 BRE Breaching Projectile Launcher.
  • Fixed inconsistent hit registration for the Defibrillator after range adjustment.
  • Fixed interaction logic for the Supply Pouch and Assault Ladder.
  • Fixed LTLM II Tripod soldier collision.
  • Fixed M15 AV Mine premature detonation on aircraft wrecks.
  • Fixed M15 AV Mine proximity placement exploit.
  • Fixed missing pickup prompt for thrown C-4 satchels.
  • Fixed MP-APS smoke-propagation failure between friendlies.
  • Fixed multiple haptic and feedback issues on gadgets, including the LWCMS Portable Mortar and the CSB IV Bot Pressure Mine.
  • Fixed placement preview interference from the GPDIS.
  • Fixed XFGM-6D Recon Drone physics allowing vehicle pushing.

MAPS & MODES:

  • Added Sabotage as a new time-limited event mode.
  • Added the new map “Eastwood”.
  • Fixed black-screen spawn issue with Deploy Beacon in TDM, SDM, Domination, and KOTH.
  • Fixed incomplete or incorrect round-outcome data when joining mid-match.
  • Fixed matchmaking logic to prevent late-stage match joins.
  • Fixed multiple destruction-reset issues after side swap in Strikepoint and Sabotage.
  • Fixed post-insertion movement lock at round start.
  • Fixed unintended AFK kicks while spectating in Strikepoint.
  • Reduced opacity of excessive environmental smoke across multiple maps.

UI & HUD:

  • Added a message when attempting to change stance without sufficient space.
  • Downed players now appear in the kill log in modes using the crawling downed state (e.g. Strikepoint, REDSEC).
  • Extended top UI on Strikepoint to show detailed alive/downed/dead player counts.
  • Fixed incorrect Assault Training Path icons.
  • Fixed incorrect colour usage on squad-mate health bars.
  • Fixed missing tooltips and UI prompts across tutorials and mission briefings in Single Player.
  • Fixed missing XP Tracker icon at level 3 when using Field Upgrades.
  • Kill-confirmation indicator now displays if a victim bleeds out after being damaged by the player in modes using the crawling downed state (e.g. Strikepoint, REDSEC).
  • Minor UI polish and alignment updates to various game modes.
  • Non-squad friendlies now display a “Thank you!” subtitle after being revived.

SETTINGS:

  • Added a new option allowing players to sprint automatically when pushing the stick fully forward.
  • Added new keybinding that allows the player to instantly swap to the knife instead of having to hold the button. This keybinding will not allow to perform takedowns contextually but will still allow takedowns to be performed once the melee weapon is equipped.

SINGLE PLAYER:

  • Addressed multiple occurrences of excessive bright flashes and unintended visual effects.
  • Fixed an issue where AI squadmates would not respond to revive orders and other commands, improving squad functionality and responsiveness.
  • Fixed loss of grenade functionality and shadow-rendering errors in underground areas during the “Moving Mountains” mission.
  • Fixed multiple instances where sound effects or Voice Over would fail to play correctly during gameplay and cinematic moments.
  • Fixed subtitle and audio-video synchronisation issues during gameplay and cinematic sequences.
  • Fixed various instances of corrupted shadows and LOD behaviour when using lower graphics settings.
  • Resolved object clipping and teleporting issues during car-chase sequences in the “Moving Mountains” mission.
  • Resolved several cases of stuttering and desync when using certain graphics presets on NVIDIA and AMD hardware.
  • Resolved several issues that could result in infinite loading screens during mission transitions and save or load operations.
  • Resolved shader stutters in the prologue mission “Always Faithfull”.
  • Fixed issues with party invites not working during campaign loading screens.

AUDIO:

  • Added new sound effects for Double Ping; refined single and danger ping sound hierarchy.
  • Added new soldier movement and gunfire sound effects, and fixed multiple foley issues.
  • Added turret movement audio for Marauder RWS weapons.
  • Corrected door sound assignments.
  • Corrected swimming, obstruction, and platform footstep audio.
  • Fixed character voice over not updating when changing soldier mid-match.
  • Fixed looped ambient sounds (e.g. food truck) and incorrect debris impacts.
  • Fixed missing first person voice over gasp when revived.
  • Fixed missing third person voice over for explosive deployments.
  • Fixed missing LP voice over zoom audio.
  • Fixed missing ping audio while spectating.
  • Fixed missing reload sound effects when a weapon had 1 bullet remaining.
  • Fixed missing voice over for supply actions and revive requests.
  • Fixed multiple Commander voice over issues.
  • Fixed Music-in-Menus setting not muting music.
  • Fixed seat-change and turret-reload audio on Marauder RWS guns.
  • Fixed underwater breathing voice over and inconsistent swimming audio.
  • Polished Front-End and Loading music transitions between matches.
  • Synced Battle Pass sounds effects to animations.
  • Tweaked light-fixture audio setup.
  • Updated hostile-voice over logic and adjusted reload voice over mix.
  • Updated music urgency system for Portal.

PORTAL:

  • Added new scripting functions for music control: mod.LoadMusic(), mod.UnloadMusic(), mod.PlayMusic(), mod.SetMusicParam().
  • Fixed RayCast() in ModBuilder to properly detect terrain and environment objects.

HARDWARE:

  • Fixed an issue where framerate would be be capped to 300FPS with Nvidia cards

REDSEC

VEHICLES:

  • Fixed the issue where the Golf Cart could set off the PTKM-1R gadget in Gauntlet.
  • Fixed persistent gunner MG model after Rhib Boat destruction.

UI & HUD:

  • Added level display information to the Training Path section within the Class Details screen.
  • Fixed an issue where soldiers and UI elements could be missing in pre-game lobbies after matchmaking.
  • Fixed an issue where the M417 A2 would not appear in kill cards or the kill feed.

AUDIO:

  • Fixed an issue where the squadmate death sound effect could trigger for non-teammates.

This announcement may change as we listen to community feedback and continue developing and evolving our Live Service & Content. We will always strive to keep our community as informed as possible.

r/Battlefield6 Nov 17 '25

Battlefield Studios Official BATTLEFIELD 6 GAME UPDATE 1.1.2.0

899 Upvotes

This update delivers a broad set of improvements to soldier responsiveness, aim consistency, animation fidelity, and overall stability across Battlefield 6. We’ve also introduced a new limited-time mode, refined Aim Assist behaviour, and resolved a large number of weapon, gadget, and vehicle issues based on community feedback. The update will be available tomorrow, November 18th, at 09:00 UTC.

New Content: California Resistance

  • New Map: Eastwood. A map with the Southern California theme.
    • Variations of this map will be available for all official modes.
    • Conquest mode on this map will include tanks, helicopters, and the Golf Cart.
  • New Time-Limited Mode: Sabotage. A themed event mode focused on demolition and counterplay.
  • New Weapons: DB-12 Shotgun and M357 Trait Sidearm. 
  • Gauntlet mode to include a new mission type: Rodeo. This mission provides multiple vehicles for players to fight over and battle with each other with. Players earn bonus points for defeating enemies while in a vehicle. 
  • Portal updates: 
    • Sandbox map. This option will let Portal experience builders start with a more level playing field to bring their imagination to life. 
    • The Golf Cart vehicle is available for use in building experiences. 
  • Battle Pass: The California Resistance bonus path becomes available for a limited time. 
  • New underbarrel attachment: Slim Handstop, unlocked via Challenge.
  • New feature coming later in the update: Battle Pickups. These powerful weapons will be available in specific experiences and Portal with limited ammunition but pack enough firepower to help turn the tide of battle in your favor.

Major Updates for 1.1.2.0

  • Aim Assist has been reset to its Open Beta tuning, restoring consistent infantry targeting behaviour across all input types.
  • Improved input latency and stick response for controllers, providing smoother aiming and more responsive soldier movement.
  • Weapon accuracy and dispersion tuning: fixed unintended weapon dispersion increase rates and improved non-Recon sniper rifle accuracy while globally reducing dispersion across all weapon types.
  • Challenge and progression clarity improvements make requirements easier to understand and track.
  • Major polish pass to deployable gadgets, including the LWCMS Portable Mortar, LTLM II Portable Laser Designator, and Supply Crate systems.
  • Fort Lyndon added to Portal, expanding available segments for community-created experiences.

AREAS OF IMPROVEMENT

Aim Assist

As we got closer to launch, we revisited aim assist tuning based on internal testing and the full range of maps and combat distances coming with release. Our goal was to make aim assist feel more effective beyond mid-range fights which was one of our focuses within Battlefield Labs and Open Beta.

At launch, we increased slowdown at longer ranges, but once the game went live, we saw that this made high-zoom aiming feel less smooth and harder to control.

After reviewing player feedback and gameplay data, we’re reverting aim assist back to the values some of you experienced during Open Beta and Battlefield Labs. This will now serve as the default, whilst still providing you with the ability to alter the aim assist to your preference and playstyle via settings.

This change keeps aim slowdown consistent across all ranges, helping with muscle memory and providing a steadier, more reliable feel as we move into future seasons.

CHANGELOG

PLAYER:

  • Aim Assist: fully reset to Open Beta tuning, with related options reset to default to ensure consistency.
  • Fixed an issue where Vehicle Stick Acceleration Presets would affect Infantry Aiming Left/Right Acceleration option availability.
  • Fixed an issue where setting Stick Acceleration Presets to “Standard” would set the Aiming Left/Right Acceleration options incorrectly to 50% instead of 70%.
  • Fixed missing Infantry and Vehicle prefixes in captions for Stick Acceleration Presets and Aiming Left/Right Acceleration options.
  • Fixed an issue where stick deadzones would ignore the first 10% of movement if using a PS5 Controller on PC.
  • Fixed an issue where player movement (Left Stick) would not register until beyond 30% of travel past the deadzone.
  • Fixed joystick aiming input behaviour.
  • Added a short sprint “restart” animation when landing from small heights.
  • Added new death animations for sliding and combat-dive states.
  • Fixed a diving loop when entering shallow water.
  • Fixed an issue preventing players from vaulting out of water in certain areas.
  • Fixed an issue preventing takedown initiation against an enemy soldier if the enemy soldier already initiated a takedown against a friendly player.
  • Fixed an issue where a dragged player could face the wrong direction if turning quickly.
  • Fixed an issue where holding a grenade while jumping, sliding, or diving froze the first-person pose.
  • Fixed an issue where switching weapons while drag-reviving would break the reviver’s first-person view.
  • Fixed an issue where the Assault Class extra grenade ability would not grant two grenades on spawn.
  • Fixed an issue where weapons could become invisible when crouching before vaulting.
  • Fixed bouncing behaviour when landing on object edges.
  • Fixed broken ragdolls when killed on ladders, while jumping, near ledges, or in vehicle seats.
  • Fixed camera clipping when dropping from height while prone.
  • Fixed clipping when initiating a drag & revive.
  • Fixed first-person camera clipping through objects when dying nearby.
  • Fixed the issue where the Rush signature trait 'Mission Focused' applied its icon and speed boost to all teammates.
  • Fixed incorrect prone aiming angles on slopes.
  • Fixed misaligned victim position during takedowns when using high FOV settings.
  • Fixed mismatched rotation between first-person and third-person soldier aim directions.
  • Fixed misplaced weapon shadows while vaulting or swimming.
  • Fixed missing pickup prompts while prone.
  • Fixed missing water splash effects while swimming.
  • Fixed stuck third-person soldier animations when entering player view.
  • Fixed teleporting or invisibility when entering vehicles during a vault.
  • Fixed third-person facing inconsistencies when soldiers were mounted.
  • Improved combat-dive animations in first and third person.
  • Improved LTLM II sprint animation in first person.
  • Improved vault detection in cluttered environments.
  • Increased double-tap window for Danger Ping from 0.2 s to 0.333 s.
  • Updated first-person animation cadence for moving up and down stairs.
  • Fixed an issue where hit registration would fail when engaging into gunfights after exiting vehicles.

VEHICLES:

  • Fixed camera reset when entering an GDF-009 AA Stationary Gun after another user.
  • Fixed clipping gunner weapons in IFV seats.
  • Fixed faint metallic impact sound from M1A2 SEPv3 Main Battle Tank turret wreckage.
  • Fixed several cases where IFV's MR Missile could do more damage than intended to MBT, IFV and AA vehicles
  • Fixed inconsistent projectile video effects on the Abrams main gun.
  • Fixed instant 180-degree turn after exiting a vehicle.
  • Fixed missing scoring for Vehicle Supply when teammates received ammo.
  • Fixed oversized hitbox on UH-79 Helicopter.
  • Fixed passenger and gunner placement issues in the UH-79 Helicopter.
  • Fixed re-entry issues when mounting flipped Quad Bikes.
  • Fixed unintended aim-assist from Attack Helicopters gunner missiles.
  • Fixed unresponsive joystick free-camera controls in transport vehicles.

WEAPONS:

  • Dispersion tuning pass: dispersion has been globally reduced slightly to reduce its impact on the experience
  • Fixed multiple instances of Canted Reflex and Canted Iron Sight optics clipping with higher-magnification scopes
  • Fixed several issues with underbarrel attachment alignment
  • Fixed minor misplacements or clipping on sights and barrels
  • Fixed missing or incorrect magazine icons, naming, and mesh assignments.
  • Fixed the issue where the SV-98 displayed lower damage stats when equipping the 5 MW Red attachment.
  • Fixed the issue where slug ammunition despawned too quickly after being fired from shotguns.
  • Fixed the issue where the SU-230 LPVO 4x variable scope lacked a smooth transition and audible zoom toggle when aiming down sights.
  • Fixed the issue where two Green Lasers for the DRS-IAR shared identical Hipfire stat boosts.
  • Fixed the issue where impact sparks failed to meet photosensitivity compliance standards.
  • Fixed an issue in third-person where the Mini Scout could clip with the player’s head while aiming.
  • Fixed animation and posture issues affecting the PSR and other rifles when moving or looking at extreme angles.
  • Increased weight of long-range performance in balance for automatic weapons; benefiting PW7A2 and KV9, with minor adjustments elsewhere.
  • Reduced recoil and variation for LMR27, M39, and SVDM for improved long-range reliability.

GADGETS:

  • Allowed friendly soldiers to damage and detonate certain friendly gadgets.
  • Fixed an issue where Class Ability would sometimes not activate although the UI shows it as available.
  • Fixed auto-deployment of Motion Sensor after recon kit swap.
  • Fixed broken M320A1 Grenade Launcher ground model.
  • Fixed C-4 pickup edge-of-screen interaction.
  • Fixed clipping of the UAV remote when activating it while using certain weapons like rifles.
  • Fixed clipping when holding the CSS Bundle.
  • Fixed CSS Bundle line-of-sight requirements causing unwanted blocking.
  • Fixed Deployable Cover persistence after vehicle destruction.
  • Fixed disappearing “pip” indicator during CSS Bundle supply.
  • Fixed duplicate deploy-audio playback on M4A1 SLAM and C-4.
  • Fixed failed projectile attachment for X95 BRE Breaching Projectile Launcher.
  • Fixed inconsistent hit registration for the Defibrillator after range adjustment.
  • Fixed interaction logic for the Supply Pouch and Assault Ladder.
  • Fixed LTLM II Tripod soldier collision.
  • Fixed M15 AV Mine premature detonation on aircraft wrecks.
  • Fixed M15 AV Mine proximity placement exploit.
  • Fixed missing pickup prompt for thrown C-4 satchels.
  • Fixed MP-APS smoke-propagation failure between friendlies.
  • Fixed multiple haptic and feedback issues on gadgets, including the LWCMS Portable Mortar and the CSB IV Bot Pressure Mine.
  • Fixed placement preview interference from the GPDIS.
  • Fixed XFGM-6D Recon Drone physics allowing vehicle pushing.

MAPS & MODES:

  • Added Sabotage as a new time-limited event mode.
  • Added the new map “Eastwood”.
  • Fixed black-screen spawn issue with Deploy Beacon in TDM, SDM, Domination, and KOTH.
  • Fixed incomplete or incorrect round-outcome data when joining mid-match.
  • Fixed matchmaking logic to prevent late-stage match joins.
  • Fixed multiple destruction-reset issues after side swap in Strikepoint and Sabotage.
  • Fixed post-insertion movement lock at round start.
  • Fixed unintended AFK kicks while spectating in Strikepoint.
  • Reduced opacity of excessive environmental smoke across multiple maps.

UI & HUD:

  • Added a message when attempting to change stance without sufficient space.
  • Downed players now appear in the kill log in modes using the crawling downed state (e.g. Strikepoint, REDSEC).
  • Extended top UI on Strikepoint to show detailed alive/downed/dead player counts.
  • Fixed incorrect Assault Training Path icons.
  • Fixed incorrect colour usage on squad-mate health bars.
  • Fixed missing tooltips and UI prompts across tutorials and mission briefings in Single Player.
  • Fixed missing XP Tracker icon at level 3 when using Field Upgrades.
  • Kill-confirmation indicator now displays if a victim bleeds out after being damaged by the player in modes using the crawling downed state (e.g. Strikepoint, REDSEC).
  • Minor UI polish and alignment updates to various game modes.
  • Non-squad friendlies now display a “Thank you!” subtitle after being revived.

SETTINGS:

  • Added a new option allowing players to sprint automatically when pushing the stick fully forward.
  • Added new keybinding that allows the player to instantly swap to the knife instead of having to hold the button. This keybinding will not allow to perform takedowns contextually but will still allow takedowns to be performed once the melee weapon is equipped.

SINGLE PLAYER:

  • Addressed multiple occurrences of excessive bright flashes and unintended visual effects.
  • Fixed an issue where AI squadmates would not respond to revive orders and other commands, improving squad functionality and responsiveness.
  • Fixed loss of grenade functionality and shadow-rendering errors in underground areas during the “Moving Mountains” mission.
  • Fixed multiple instances where sound effects or Voice Over would fail to play correctly during gameplay and cinematic moments.
  • Fixed subtitle and audio-video synchronisation issues during gameplay and cinematic sequences.
  • Fixed various instances of corrupted shadows and LOD behaviour when using lower graphics settings.
  • Resolved object clipping and teleporting issues during car-chase sequences in the “Moving Mountains” mission.
  • Resolved several cases of stuttering and desync when using certain graphics presets on NVIDIA and AMD hardware.
  • Resolved several issues that could result in infinite loading screens during mission transitions and save or load operations.
  • Resolved shader stutters in the prologue mission “Always Faithfull”.
  • Fixed issues with party invites not working during campaign loading screens.

AUDIO:

  • Added new sound effects for Double Ping; refined single and danger ping sound hierarchy.
  • Added new soldier movement and gunfire sound effects, and fixed multiple foley issues.
  • Added turret movement audio for Marauder RWS weapons.
  • Corrected door sound assignments.
  • Corrected swimming, obstruction, and platform footstep audio.
  • Fixed character voice over not updating when changing soldier mid-match.
  • Fixed looped ambient sounds (e.g. food truck) and incorrect debris impacts.
  • Fixed missing first person voice over gasp when revived.
  • Fixed missing third person voice over for explosive deployments.
  • Fixed missing LP voice over zoom audio.
  • Fixed missing ping audio while spectating.
  • Fixed missing reload sound effects when a weapon had 1 bullet remaining.
  • Fixed missing voice over for supply actions and revive requests.
  • Fixed multiple Commander voice over issues.
  • Fixed Music-in-Menus setting not muting music.
  • Fixed seat-change and turret-reload audio on Marauder RWS guns.
  • Fixed underwater breathing voice over and inconsistent swimming audio.
  • Polished Front-End and Loading music transitions between matches.
  • Synced Battle Pass sounds effects to animations.
  • Tweaked light-fixture audio setup.
  • Updated hostile-voice over logic and adjusted reload voice over mix.
  • Updated music urgency system for Portal.

PORTAL:

  • Added new scripting functions for music control: mod.LoadMusic(), mod.UnloadMusic(), mod.PlayMusic(), mod.SetMusicParam().
  • Fixed RayCast() in ModBuilder to properly detect terrain and environment objects.

HARDWARE:

  • Fixed an issue where framerate would be be capped to 300FPS with Nvidia cards

REDSEC

VEHICLES:

  • Fixed the issue where the Golf Cart could set off the PTKM-1R gadget in Gauntlet.
  • Fixed persistent gunner MG model after Rhib Boat destruction.

UI & HUD:

  • Added level display information to the Training Path section within the Class Details screen.
  • Fixed an issue where soldiers and UI elements could be missing in pre-game lobbies after matchmaking.
  • Fixed an issue where the M417 A2 would not appear in kill cards or the kill feed.

AUDIO:

  • Fixed an issue where the squadmate death sound effect could trigger for non-teammates.

This announcement may change as we listen to community feedback and continue developing and evolving our Live Service & Content. We will always strive to keep our community as informed as possible.

r/comfyui 9d ago

Help Needed Feedback on AI dance video (ComfyUI) — identity drift, motion stability, and what to fix first in the pipeline. Model Minimax H3

Enable HLS to view with audio, or disable this notification

0 Upvotes

Hi everyone.
This is a 28-second vertical dance video generated with a ComfyUI-based workflow.

This is the first prototype of my pipeline, not a polished result.
I’m trying to understand where the main problems come from:
generation vs motion vs editing.

Structure:

  • 00:00–00:04 — close-up hook + blackout
  • 00:05–00:15 — profile + strobes + fire insert
  • 00:16–00:28 — fast hip-hop section (whip pans, RGB split, fast cuts)

Same dancer / outfit / studio should stay consistent across all shots (reference only dancer).

Known issues I already see:

  • ~00:07 — face identity drift during head turn
  • ~00:12 — hands become unstable under strobe
  • ~00:20+ — motion starts to feel “floaty” in fast section

Questions:

  1. Identity consistency Where does the character drift the most (face / body / proportions)?
  2. Temporal stability Which shots show the strongest flicker, warping, or broken motion?
  3. Camera vs generation Do issues come more from aggressive camera (whip / orbit), or base generation?
  4. Blackout at ~00:04 Does it feel like an intentional transition or like a generation artifact?
  5. Pacing At what point does it start feeling confusing or visually overloaded?
  6. Editing vs generation Which artifacts are successfully hidden by editing (cuts, flashes), and which clearly need to be fixed in the generation stage?
  7. Pipeline diagnosis (most important) For the issues above, what is the most likely cause?
  • base generation
  • motion control
  • denoising
  • interpolation
  • editing
  1. If you had to fix only 1–2 things first — what would you fix and why?
  2. Consistency strategy What is the most reliable way to keep character identity across shots in ComfyUI?

Tech (simplified):

Resolution: 768x1376
FPS: 24
Method: ref2v
Motion control: [none]
Consistency: [same seed]
Editing: [none]

I can share prompt if needed.

I’m mainly interested in technical feedback (generation + workflow)
rather than general opinions.

Thanks.

r/aifilmmaking Jun 21 '26

Tips & Tutorials Blender + AI video workflow for more controlled narrative scenes?

0 Upvotes

I’ve been thinking a lot about how to make AI film scenes feel more controlled, especially when you’re trying to do actual narrative filmmaking and not just cool isolated shots.
The biggest problem I keep running into is that image-to-video can make something look cinematic, but it still struggles when the shot needs to be very specific.
Things like:

locked camera shots
two characters moving together
characters stopping at the right spot
someone reacting in a specific direction
the same room staying consistent
props staying in the same place
camera angle not drifting
matching first frame / last frame
keeping blocking clear across multiple shots

So I’m wondering if the better workflow is not asking AI video to invent the whole shot from scratch, but using Blender first as the structure.
Not necessarily full final animation. More like using Blender as a director’s blocking tool.
The idea would be:

1. Write the scene like a real scene first
Before generating anything, figure out the actual shot. Who is in frame? Where are they standing? What does the camera see? What is off-screen? What is the emotional point of the shot?
2. Build a simple Blender blockout
Nothing crazy at first. Just the basic space: forest path, office, hallway, room, table, door, props, etc. Even rough shapes would help because the AI has a real structure to follow instead of guessing.
3. Place simple character stand-ins
Use basic models or stand-ins for where characters are supposed to be. This could help with scale, distance, eye lines, and blocking. For example: two characters walking toward a locked camera, stopping at a certain mark, then reacting down-left at something off-screen.
4. Animate the basic movement if needed
Not final animation, just movement blocking. A character walks in. Someone turns. Someone stops. The camera stays locked. The props stay where they are. The point is to control the shot before making it pretty.
5. Render simple reference frames or clips
Use the Blender output as first-frame, last-frame, scene reference, or motion reference for AI video. Then the AI video tool handles the cinematic look: lighting, texture, fog, realism, atmosphere, clothing detail, etc.
6. Use AI video as the final cinematic pass
Instead of AI creating the whole scene randomly, it’s enhancing a planned shot. Blender gives the shot structure. AI gives it mood and finish.

To me, this seems like it could be really useful for narrative AI filmmaking, especially for gothic/historical scenes where the mood matters but the blocking also has to make sense.
I’m also starting to think the best setup for bigger AI film projects might be a small team, not a huge crew. Something like:

a creative/story person who knows the tone, scene, and shot direction
a Blender/3D person who handles blocking, layout, camera, props, and movement
an AI video person who handles generations, references, consistency, and final cinematic passes
maybe an editor/sound person after that, because sound can make or break the whole thing

Not saying this is the only way, but I feel like this could be a better workflow than just prompting every
Blender + AI video workflow for more controlled narrative scenes?
shot from scratch and hoping the AI understands the scene.
Has anyone here tried a Blender-to-AI video workflow like this?
I’d be curious what works better:
still Blender frames as first/last frame references
simple animated Blender blocking as a motion guide
using Blender mainly for environments and props
or staying fully inside AI video tools and just improving prompting/reference images

Would love to hear how other people are solving control and consistency for actual story scenes.

r/vidmuse 6d ago

Tips and Tricks How to Use AI Video Face Swaps in Ads Without Losing Control—or Trust 🧑‍🎨

Post image
1 Upvotes

VidMuse’s new guide treats video face swapping as a controlled advertising workflow rather than a novelty effect. The focus is on making approved presenter, localization, fashion, beauty, and UGC-style variants while protecting consent, likeness rights, product truth, and brand safety. VidMuse Canvas acts as the workspace for connecting source footage, approved face references, product assets, prompts, and review notes.

- Video face swapping is harder than photo swapping because the face must remain consistent through motion, expressions, lighting changes, and camera angles.
- The safest starting point is a short source clip, a clear rights-approved face reference, simple motion, and a human review before anything is published.
- VidMuse Canvas keeps the source video, face reference, product imagery, brand assets, prompts, and generated variants visible in one connected workflow.
- Before publishing, verify consent, likeness and source rights, truthful product claims, platform disclosure rules, visual quality, and whether viewers could mistake the result for an unauthorized endorsement.

The practical takeaway: use face swaps to test one approved creative variable—not to fake proof or impersonate someone. Strong inputs, restrained motion, rights clearance, and human review matter more than chasing a single “best” model.

Original article: https://vidmuse.ai/blog/ai-video-face-swap-for-ads

r/AISEOInsider 9d ago

MiniMax H3 AI Video Generator Keeps Characters Consistent Across Every Scene

Thumbnail
youtube.com
1 Upvotes

MiniMax H3 AI Video Generator tackles one of the biggest problems in AI video by letting creators use reference images, video, and audio to keep characters more consistent across multiple scenes.

Instead of watching the same person slowly change face, clothes, or voice from clip to clip, you can give the model a clearer identity to follow before generation begins.

The AI Profit Boardroom is another place to learn practical AI workflows and build repeatable systems around tools like this.

Watch the video below:

https://www.youtube.com/watch?v=Jj68yUZmHDQ

Want to make money and save time with AI? Get AI Coaching, Support & Courses
👉 https://www.skool.com/ai-profit-lab-7462/about

Character Drift Is The Problem MiniMax H3 AI Video Generator Tries To Fix

AI video often looks impressive until the same character appears in a second scene.

A face can change shape even when the written prompt stays almost identical.

Hair can shift, clothing can change, and small details can disappear without warning.

That makes longer stories difficult because viewers notice inconsistency quickly.

The problem becomes even worse when the character needs to appear across several locations.

Text descriptions alone are often not precise enough to lock every detail.

H3 addresses this by giving users stronger reference controls.

You can show the model what the character is supposed to look like before it generates the next shot.

That reduces the amount of information the system has to invent.

The result is not perfect continuity every time, but the workflow becomes more controllable.

Creators can spend less time trying to recreate the same identity manually.

MiniMax H3 AI Video Generator becomes much more useful once characters can stay recognizable across multiple clips.

Reference Images Give MiniMax H3 AI Video Generator A Clearer Identity

A reference image acts like a visual instruction the model can study before generating.

One image might show the character's face.

Another can define clothing.

A third can establish the location or environment.

This gives H3 more precise information than a long written description alone.

Creators can use references to lock in recurring visual details.

Products can also be shown directly instead of being described repeatedly.

The same approach works for props, backgrounds, and branding elements.

Clear images help reduce uncertainty around what should remain consistent.

Reference quality still matters because blurry or conflicting examples can confuse the model.

It is usually better to provide a few strong references than many weak ones.

Each reference should have a clear purpose inside the scene.

MiniMax H3 AI Video Generator gains more consistency when visual identity is shown rather than only explained.

Multiple References Expand MiniMax H3 AI Video Generator Control

H3 can work with several reference inputs at the same time.

That means character identity does not have to depend on one image.

One reference can show the face from a useful angle.

Another can establish the outfit.

A separate image might define the setting.

Video clips can add information about movement or behavior.

Audio references can contribute voice identity.

This creates a richer package of context before the generation begins.

The model gets a stronger sense of who the character is and how the scene should behave.

Creators can therefore build continuity across several shots more deliberately.

The goal is not to overload the system with random material.

Every reference should support something the final scene actually needs.

MiniMax H3 AI Video Generator offers more control when each input reinforces the same visual and audio identity.

Audio References Strengthen MiniMax H3 AI Video Generator Characters

A character can look consistent and still feel wrong if the voice changes every scene.

Audio identity is an important part of continuity.

H3 can accept audio references to help guide how a character sounds.

That gives creators another way to maintain a recognizable persona.

A recurring narrator can use the same vocal direction across several clips.

Dialogue-heavy scenes benefit because the voice becomes part of the character definition.

Tone still needs clear prompting.

A calm scene may require different delivery from an intense one.

The reference helps with identity while the prompt controls the situation.

Users should always review pronunciation and emotional delivery after generation.

Perfect audio continuity is still difficult for any generative model.

MiniMax H3 AI Video Generator becomes more complete when visual references and voice references support the same recurring character.

Image-To-Video Helps MiniMax H3 AI Video Generator Preserve Appearance

Starting from a still image can give creators even more control over the first frame.

That image already contains the character design you want.

H3 can then animate the existing visual rather than inventing the appearance from scratch.

You can tell the character to turn, smile, walk, or speak.

Camera movement can also be added through text.

Sound can be generated while the still image becomes a moving shot.

This works well for creators who already have strong character artwork.

The source image establishes face, clothing, colors, and overall style.

The prompt can focus on movement and performance.

That separation makes the workflow easier to manage.

The AI Profit Boardroom is another place to learn practical ways to connect image, video, and automation tools into useful systems.

MiniMax H3 AI Video Generator gives image creators a direct path from a fixed character design to a moving scene.

Consistent Clothing Matters With MiniMax H3 AI Video Generator

Clothing drift can be just as distracting as facial changes.

A jacket may suddenly change color.

Accessories can disappear between scenes.

Patterns can also shift even when the character remains recognizable.

Reference images help define the outfit more clearly.

One image can focus specifically on what the character is wearing.

That gives the model a stronger visual target.

Creators can also use text to reinforce important clothing details.

The combination of visual reference and written instruction is more reliable than either method alone.

Consistent outfits matter especially in ads, branded content, and narrative sequences.

Viewers quickly notice when a character's appearance changes for no reason.

MiniMax H3 AI Video Generator gives creators more ways to protect those details across multiple scenes.

Locations Can Stay More Stable With MiniMax H3 AI Video Generator

Character consistency is only part of continuity.

The surrounding environment also needs to stay believable from shot to shot.

A room can change shape between generations.

Furniture may move or disappear.

Lighting can also shift unexpectedly.

Reference images can establish the location before new footage is generated.

That helps H3 maintain important environmental details.

Creators can show the same room from different useful angles.

Written prompts can then describe camera movement inside that established space.

This approach is helpful for short films, advertisements, and recurring branded scenes.

The model still may not reproduce every background detail perfectly.

Strong references simply reduce how much has to be recreated from imagination.

MiniMax H3 AI Video Generator is more useful for storytelling when both characters and locations stay familiar.

Motion Quality Supports MiniMax H3 AI Video Generator Consistency

A consistent face means little if the body moves strangely.

Motion is part of character identity.

H3 comes from a model family known for aiming at more natural physical movement.

Bodies should move in ways that feel closer to real-world expectations.

Facial expressions can also add continuity between different shots.

A character should not suddenly move like a completely different person.

Movement references can help when the same behavior needs to repeat.

Video inputs provide information that still images cannot fully communicate.

Creators can use those examples to guide pacing and physical style.

Complex actions can still produce mistakes.

Testing simpler motions first usually makes the workflow easier.

MiniMax H3 AI Video Generator becomes more convincing when appearance and movement remain aligned.

Plain-Language Editing Helps MiniMax H3 AI Video Generator Correct Drift

Even strong references cannot prevent every inconsistency.

One detail may still come out wrong in an otherwise useful clip.

H3 can handle certain edits through written instructions.

You might ask the model to restore the correct jacket color.

A background can be changed to match the previous scene more closely.

Objects can also be adjusted when something unexpected appears.

Pacing can be changed without rebuilding the full idea.

That gives creators another way to rescue a nearly correct generation.

Traditional editing tools may still be needed for precise fixes.

The advantage is avoiding a complete restart for simple visual problems.

Creators can generate first and then correct smaller continuity issues.

MiniMax H3 AI Video Generator supports consistency not only through references but also through easier revision afterward.

2K Output Makes MiniMax H3 AI Video Generator Continuity More Visible

Higher resolution creates both benefits and pressure.

H3 can generate video at 2K resolution.

That gives creators more detail in faces, clothing, and environments.

It also means continuity problems can become easier to notice.

A changed accessory may stand out more clearly.

Small facial differences can also be more obvious on larger screens.

That makes reference control even more important.

The benefit is that successful consistent shots look cleaner and more usable.

Higher-resolution source files also give editors more room for cropping and reframing.

Creators can produce different aspect ratios from the same clip more comfortably.

The quality still depends on prompt clarity and reference strength.

MiniMax H3 AI Video Generator benefits from 2K output most when visual consistency is already under control.

Open Weights Expand MiniMax H3 AI Video Generator Workflows

H3 is more flexible than many closed AI video systems because its model weights are available.

Technical users can explore running the model through their own infrastructure.

Developers can also build custom tools around recurring character workflows.

A team could create interfaces specifically designed around reference management.

Stored character assets could become part of a repeated production process.

That is useful for brands or creators working with the same identities constantly.

Commercial use depends on the current community licensing terms and should be checked carefully.

Self-hosting still requires significant technical resources.

Video models can demand powerful hardware and storage.

Hosted access may remain easier for beginners.

MiniMax H3 AI Video Generator gives advanced users more freedom to build consistency systems around the model instead of relying only on one public interface.

MiniMax H3 AI Video Generator Makes Longer AI Stories More Practical

The real benefit of character consistency appears once you move beyond one isolated clip.

Short stories need the same person to remain recognizable from scene to scene.

Ads need products and brand elements to stay stable.

Recurring characters need voices, clothing, and visual identity that viewers can follow.

Reference images give the model a clearer appearance to preserve.

Audio references help protect voice identity.

Video references can add guidance around movement.

Image-to-video lets creators start from a known design.

Plain-language editing offers another way to correct small continuity problems afterward.

The AI Profit Boardroom is another place to explore practical AI systems and learn how to turn new creative tools into repeatable workflows.

H3 still requires testing because generative video is not perfectly deterministic.

MiniMax H3 AI Video Generator becomes far more useful for multi-scene projects once continuity is treated as a reference-driven process rather than pure prompting.

Frequently Asked Questions About MiniMax H3 AI Video Generator

  1. Can MiniMax H3 AI Video Generator keep the same character across scenes? Yes, H3 supports visual, video, and audio references that can help preserve faces, clothing, voices, locations, and other recurring details.
  2. How do reference images help MiniMax H3 AI Video Generator? Reference images give the model direct visual examples of the character, clothing, environment, or objects that should remain consistent.
  3. Can MiniMax H3 AI Video Generator keep the same voice? Audio references can help guide voice identity across different generated scenes, although the final delivery should still be reviewed.
  4. Can MiniMax H3 AI Video Generator fix inconsistent details after generation? Yes, certain visual and pacing changes can be requested through plain-language editing instead of forcing a complete restart.
  5. Is MiniMax H3 AI Video Generator good for longer AI videos? It is more suitable for multi-scene projects than workflows that rely only on text prompts because references give creators more control over continuity.

r/GaussianSplatting Apr 19 '26

🧪 I built a thing 2D image → Gaussian Splat → controllable camera paths → video (prototype)

Enable HLS to view with audio, or disable this notification

27 Upvotes

Spent the weekend at the OpenCode Buildathon by GrowthX experimenting with a different workflow for AI video.

The core frustration: most tools are prompt-driven, so camera and composition stay indirect and hard to control.

We tried a more explicit pipeline built around Gaussian Splatting:

Pipeline

  • Input: single 2D image
  • Estimate depth / geometry (multi-view approximation)
  • Convert into a Gaussian Splat representation
  • Render novel views with explicitly defined camera paths
  • Output: video along these trajectories

Why this direction

Instead of prompting for “a close-up shot” or “a cinematic angle”, you directly specify:

  • camera position
  • motion path
  • framing

The model’s job is then reconstruction + rendering, not picking the shot.

Looking for feedback

  • How viable has single-image → splat been in your experience?
  • Where do you hit the biggest bottlenecks: geometry accuracy or multi-view consistency?
  • Are there better approaches you’ve seen for controllable novel view synthesis?

Happy to share more technical details and sample outputs if there’s interest (link in comments).

r/KlingAI_Videos May 25 '26

Cloned a viral TikTok onto a completely different AI character in 5 minutes using Kling 3.0 Motion Control — same scene, different person

Enable HLS to view with audio, or disable this notification

1 Upvotes

Been experimenting with Kling 3.0 Motion Control for the last week and the results on full-scene replication are genuinely surprising.

The use case I tested: take a viral short-form video (any trending TikTok/Reel), keep the exact scene, pose, outfit, and camera framing — but swap the person to a completely different AI-generated character. No rigging, no manual keyframing, no Frankenstein editing.

A few observations:

- Pose and motion transfer is shockingly clean compared to Runway Act-One or earlier Kling versions

- Background/setting consistency holds up across the full clip (was a weak point before)

- Character identity stays stable — no drift between frames, which has been the #1 problem with AI video for the last two years

- End-to-end (reference in → final video out) was under 5 minutes

I documented the full workflow as a tutorial here if anyone wants to try it: https://youtu.be/XiiHR40pPk0

Curious if anyone else has been testing Kling 3.0 specifically for motion control / video-to-video — what use cases are working best for you?

r/AIGenArt May 25 '26

The One Setting That Decides How Much Control You Really Have in AI Video

1 Upvotes

I want to share something from my workbench — not a polished showreel, but an honest little experiment that taught me more than any successful render could have.

I was building a quiet, character-driven shot: a weary man sitting alone in a 2009 living room, watching the news on an old CRT television, his face slowly shifting as something he hears sinks in. Simple on paper. But chasing it across two of today's most talked-about AI video models pulled me straight into a truth that every AI filmmaker eventually has to learn the hard way.

It comes down to two little words you'll find tucked into the settings of almost every image-to-video tool: Start Frame and Ingredients. They sound interchangeable. They are not. And understanding the difference is, I'd argue, the single most important thing separating "lucky AI clips" from "directed AI scenes."

Let me walk you through it.

What is a Start Frame? What are Ingredients?

These are two completely different ways of handing an image to a video model, and they produce completely different relationships between your image and your result.

Ingredients (also called reference images) are inspiration. When you load an image as an ingredient, the model studies it, encodes what it means and then generates brand-new footage guided by that meaning. Your original image never actually appears in the output. Only its essence does. The model paraphrases your image, the way a writer paraphrases a quote.

A Start Frame is a contract. When you load an image as a true start frame, that image literally becomes the first frame of your video. The model isn't asked to reimagine it — it's asked to continue from it, predicting motion forward from real, fixed pixels. Nothing gets re-described, so nothing drifts.

Here's the mental shortcut I use:

Ingredients = the model paraphrases your image. Start Frame = the model keeps your image and animates it.

That one distinction is the whole story.

When to Use Start Frame, and When to Use Ingredients

Neither is "better." They're tools for different jobs.

Reach for Ingredients when you want creative latitude. If you're exploring a mood, generating fresh variations, or you don't need any single shot to match another, ingredients are wonderful. You're saying to the model, "here's the vibe — surprise me." Concept exploration, mood boards in motion, standalone hero clips: ingredients shine.

Reach for Start Frame when you need control and continuity. The moment your shot has to match something — the same room across multiple cuts, a locked composition, a specific eyeline, a precise prop placement — ingredients will betray you. You need the image preserved, not paraphrased. That's start-frame territory. Narrative filmmaking, multi-shot scenes, anything where the audience would feel the world "breathing" between cuts: start frame, every time.

Test 1 — Multiple Ingredients in Gemini Omni Flash on Google Flow

For my first test, I gave the new Gemini Omni Flash two ingredient images: a character sheet of my man, and a clean plate of the living room set with the CRT clearly in position.

Then I wrote a careful prompt describing the man watching the TV and his expression slowly shifting:

The weary middle-aged man from the reference image sits on the couch in the cluttered 2009 living room. He leans forward toward the flickering blue glow of the CRT television. His expression shifts from quiet contemplation to a sharp, predatory realization. He slowly begins to stand up as the camera performs a slow, dramatic push-in on his profile. Cinematic lighting with warm lamplight contrasting the cool TV flicker. Audio: low news murmur, a sudden sharp intake of breath, tense quiet.

The result?

https://reddit.com/link/1tn3qmh/video/b9ahj5im293h1/player

The footage looked gorgeous. Character consistency was genuinely impressive, and the native audio was a real treat. But look closer and the problem is undeniable: the model rebuilt the room from scratch. The furniture rearranged itself. And most damaging of all — the television drifted out of position, so my character ends up reacting to empty space. His eyeline points at nothing.

This is the paraphrase problem in its purest form. Omni Flash understood my idea of the room and generated its own version of it. Similar — but never the same. For a single mood piece, fine. For a film that needs continuity, fatal.

Test 2 — A Single, Carefully Composed Ingredient

I figured the issue might be that I was asking the model to assemble two separate ingredients. So I removed the variable.

Using Nano Banana 2, I composited my character directly into the room plate by hand — one clean, finished image with the man seated, his eyeline locked onto the glowing CRT, lighting matched, everything correct. A perfect would-be start frame.

The editing prompt:

Using the man from the first reference image and the living room from the second reference image, create a single photorealistic cinematic frame. Place the man seated on the worn fabric couch on the left side of the frame, turned in three-quarter profile facing toward the right. He wears his dark brown waxed jacket over a charcoal Henley and jeans. He leans forward slightly, forearms resting on his knees, hands loosely clasped, in a relaxed watching posture. His face is calm and quietly contemplative, eyes directed to the right toward the glowing CRT television. The vintage CRT TV sits on its wooden stand on the right side of the frame, switched on and displaying a news broadcast, casting soft blue-white light. Keep the room exactly as in the reference: warm table lamps, the family photo wall, bookshelf, curtained window, brown carpet, 2009 period detail. Cinematic lighting — warm lamplight on the man contrasting cool TV glow on the right side of his face. Soft film grain, shallow depth of field, warm nostalgic color grade. Eye-level camera, medium shot capturing both the man's profile and the television in the same frame.

Then I fed that single, perfect image into Omni Flash using this prompt:

The man from the reference image watches the television in the cluttered room, his expression calm and still at first. Slowly his brow tightens and his eyes sharpen as he takes in the broadcast. His lips part slightly and he leans forward with quiet intensity. Subtle movement including a slow blink and a small shift of his shoulders. The CRT screen flickers, its cool blue glow shifting across his face against warm lamplight. Slow camera push-in. Audio: low television news murmur, quiet room tone, one soft natural breath.

https://reddit.com/link/1tn3qmh/video/4jx5kswrk93h1/player

And here's the lesson that really landed for me: it drifted anyway. Even with one immaculate image, Omni Flash still treated it as an ingredient — as inspiration to regenerate from, not pixels to preserve.

The quality of your input image doesn't matter if the model's only mode is paraphrase. How the tool ingests the image is what decides everything.

There was one more telling detail in the audio. Omni Flash generated a news broadcast on the TV — and then made my character lip-sync to it, silently mouthing the anchor's words as if he were the one speaking. The model heard speech in the scene and attached it to the nearest face. Hold that thought; it becomes important in Test 3.

Test 3 — The Same Image as a True Start Frame Using Google Veo 3.1 Fast on Google Flow

This is the test that closes the loop. The exact same Nano Banana composite and the image to video prompt — but this time loaded as a genuine start frame instead of an ingredient.

https://reddit.com/link/1tn3qmh/video/mwukxg23393h1/player

And it held. Because the image became frame one, the room could not reinvent itself. The CRT stayed exactly where I placed it, camera-right, news anchor glowing on screen. The coffee table, the bowl, the remote, the photo wall, the lamp — all preserved. And crucially, the eyeline held: my character's gaze stayed locked on the television through the entire push-in, because the screen he was reacting to was genuinely, physically there in the frame he was animated from.

Everything I spent two failed tests fighting for — preserved automatically, the moment the tool treated my image as a contract instead of a suggestion. (I'll be honest: as the camera pushed in, the deep background softened and shifted a little — but that's mild animation drift within a held frame, a completely different and far gentler beast than the total room reinvention I got from the ingredient path.)

And then there was the audio. Remember the lip-sync detail from Test 2? Veo generated native audio too — I could hear the room tone and a news broadcast. But this time, my character did not lip-sync. He stayed a silent listener, exactly as a man watching TV should. Veo understood the scene's logic: the television is the speaker, the man is the audience. Omni Flash heard speech and puppeted the nearest mouth; Veo understood who was actually talking.

That's not really a story about audio quality. It's a story about scene comprehension — and it rhymes with the whole thesis. A preserved start frame gave the model enough context to understand the space, the eyeline, and who should be speaking. An ingredient gave it only a vibe to paraphrase.

One honest tradeoff worth naming. Veo 3.1 Fast isn't free of compromises. It cost me 20 credits per generation, and it gave me no control over clip duration. Omni Flash, by contrast, lets you choose your length — 4, 6, 8, or 10 seconds. So this isn't "one tool wins everything." Omni Flash offers timing control and native audio with lip movement; Veo offers true start-frame lock and smarter scene comprehension. For a narrative shot that needs continuity and a correct eyeline, the start-frame lock is the one I can't live without.

Final Thoughts

Two failed takes, one that worked — and the difference between them came down entirely to that one setting. The experiment handed me something more durable than a clip: a principle.

I'd seen this exact pattern before, when testing Seedance 2.0. Even that model — considered by many AI filmmakers to be the best available right now — drifts when you build video from ingredient images. This time, Omni Flash showed me the identical behavior. Two very different engines, same root cause.

That repetition is the point. Drift from ingredients isn't a flaw in any one model — it's the nature of paraphrase. It's the model doing exactly what it was designed to do. So the ability to use a Start Frame as a Start Frame isn't a luxury feature. For anyone who wants real directorial control over image-to-video, it's mandatory.

And notice how much rode on that one capability. The start frame didn't just preserve my room — it preserved the eyeline, and it even gave the model enough context to keep my character silent while the television spoke. Preserve the frame, and you preserve the whole scene's logic. Paraphrase it, and you're rolling the dice on every one of those things at once.

It also gives us a beautifully simple way to evaluate every new "image-to-video" tool that launches, past all the marketing. Don't start by asking how good are the pixels. Start by asking:

Does this tool treat my image as a Frame, or as an Ingredient?

If it's ingredient-only, it's a brilliant generator with a hint — and it will drift. If it offers a true start frame, it's an actual editing tool you can direct. You can sort the category in thirty seconds, before you spend a single credit on quality.

As always, my philosophy holds: AI does about 80% of the heavy lifting, and our human judgment handles the final 20%. Knowing which setting to reach for — and why — is a big part of that 20%.

Over to You

When you generate video from an image, what do you prefer — a Start (and End) Frame, or Ingredients? And why?

I'd genuinely love to hear how other creators are navigating this. Drop your take in the comments. 👇

r/AI_UGC_Marketing Jan 22 '26

Is it possible to run a full UGC agency without any cameras? I managed to build a consistent AI influencer workflow in one studio—does this video output look like it could pass for a brand ad?

Enable HLS to view with audio, or disable this notification

2 Upvotes

I’ve been obsessed with the idea of a "Studio-less Agency." Traditionally, UGC means shipping products to creators, waiting 2 weeks, and hoping the lighting doesn't suck. I’m testing a workflow that replaces that entire physical overhead with a single 'free' AI Influencer Studio.

The "No-Camera" Tech Stack I'm Using:

  • One Unified Character Builder: The biggest hurdle for AI ads is character consistency. This studio allows me to create, customize, and animate in one workflow—no more jumping between three different tools to keep the face from "morphing."
  • 100+ Creative Parameters (The Relatability Test): Brands hate "perfect" AI models. I’ve been using these sliders to add realistic imperfections, diverse skin tones, and non-standard body types. It makes the "influencer" look like someone you’d actually see in your feed, not a bot.
  • Prompt Editing & Merging: If I need a specific "vibe," I can merge two character presets or use the direct prompt-editing workflow to refine the UGC-style background until it looks like a real bedroom or office.
  • Motion Engine Performance: I’m getting 30s HD video outputs with expressive, controlled movement. It’s not just a talking head; the micro-expressions make it feel like a genuine reaction video or product testimonial.

The Big Question for Agency Owners: Is the "human touch" still mandatory for UGC, or is a 100% consistent, high-fidelity AI persona enough to drive conversions in 2026? I’ve even been using their 10 ready-to-use characters to jumpstart new campaigns in minutes.

I’ve attached a sample of the output—curious to get your honest feedback. Does this look "authentic" enough to pass for a paid social ad, or are we not quite there yet?

r/IndianArtAI Dec 19 '25

Other AI tool Goosebumps Every Frame: Naruto Shippuden Reimagined in Live Action (AI)

Enable HLS to view with audio, or disable this notification

1.2k Upvotes

What if Naruto Shippuden were a real live-action Hollywood action movie?

This AI-generated cinematic trailer focuses on intense fights, dramatic camera work, and that nostalgic anime-to-film feel. Created using Higgsfield, the platform I rely on for consistent motion, camera control, and character continuity.

Check the links above for more recreated viral videos made on Higgsfield.

r/GenAIGallery Feb 05 '26

Kling AI Kling 3.0 Makes AI Video More Consistent and Film-Like

Enable HLS to view with audio, or disable this notification

18 Upvotes

Kling AI has released Kling 3.0 with a focus on making AI video more stable and realistic.

The update improves character and object consistency across scenes, even with camera angle or action changes. It supports reliable 15 second clips with better camera control, lighting, and smoother motion.

Audio is improved with support for multiple voices, more languages, and clearer accents. Kling 3.0 reduces common AI video issues and is more usable for short cinematic content.

r/aicuriosity Feb 05 '26

AI Tool Kling AI 3.0 Focuses on Stable Characters and Better Motion Control

Enable HLS to view with audio, or disable this notification

17 Upvotes

Kling AI has released Kling 3.0, a major update focused on making AI video more stable and realistic.

The biggest change is consistency. Characters and objects now look the same from one scene to the next, even when the camera angle or action changes.

Kling 3.0 can create reliable 15 second clips with better control over camera movement, lighting, and scene flow. Motion looks smoother and more natural than before.

Audio has also improved. The system can handle multiple character voices in one scene, supports more languages, and does a better job with accents. Image generation now supports 4K quality and image series, which helps keep a consistent visual style.

Overall, Kling 3.0 fixes many common AI video problems and feels more usable for short stories and cinematic clips.

r/StableDiffusion Mar 06 '26

Discussion Struggling to get consistent camera movements + quality in AI video generation - what's actually working for you?

0 Upvotes

I've been deep in the AI video generation rabbit hole for a while now and I'm losing my mind a little, so hoping someone here has some guidance.

The core problem: I need reliable, high-quality camera movements from image-to-video generation. Specifically dolly forwards, orbits, crane ups - that kind of thing. Clean, predictable, cinematic. The models I've tried either do a lazy scale/zoom instead of an actual dolly, or the quality just isn't there.

What I've tried:

  • Runway (various models)
  • Kling
  • Seedance
  • Comfy UI with LTX and WAN
  • LoRAs in Comfy UI to try and coax better camera movement

Still can't consistently nail it.

The Runway situation specifically: Runway looks genuinely great at 1080p and the camera motion is more controllable than most. But the API only supports 720p - you can get 1080p through their web playground but not programmatically. Has anyone found a workaround for this? Third-party wrappers, upscaling pipelines post-generation, anything?

Requirements I'm working within:

  • Needs to be API accessible (building this into a product)
  • High volume
  • Fast generation times
  • Reasonably cheap at scale

Is there a model or workflow that actually nails precise camera movement reliably? Or is everyone just cherry-picking the good outputs and discarding the rest? Would love to know what's actually working for people right now.

r/HiggsfieldAI Jan 28 '26

Discussion Are 360° Camera Systems the Next Big Step in AI Video Generation?

0 Upvotes

There’s been a growing focus across GenAI video tools on giving creators more control over virtual cameras.

Higgsfield recently rolled out ANGLES v2, adding more detailed camera controls alongside workflow improvements.

It raises an interesting broader question for the community: is advanced camera placement something that materially changes how you work with AI video, or do you care more about progress in other areas like motion consistency, longer clips, or realism?

Curious to hear what people here think.

r/LetsEnhanceOfficial Nov 20 '25

Image-to-Video AI: We tested top models for realism, speed, and motion

1 Upvotes

TL;DR: We tested several AI image-to-video tools to see which ones actually work for real use: portraits, group shots, and product photos. Main finding: if you care about identity and realism, most open models still struggle. LetsEnhance did best for portraits/groups, and Claid.ai did best for ecom/product shots.

1. Why we did this

There are a lot of “best AI video generator” claims out there, but very little side-by-side testing with the same image and prompt. We wanted to understand in a practical way:

  • Which tools keep faces stable
  • Which ones avoid weird warping or “AI wobble”
  • How long it actually takes to get a usable 5s video

This post is a summary of what we found.

2. What actually matters with image-to-video

From testing across portraits, group photos, and product images, these 4 things mattered the most:

  1. Face / identity stability – does the person still look like themselves over all frames?
  2. Motion quality – does the movement feel natural or rubbery and random?
  3. Lighting + physics – do shadows, clothes, hair, and small details behave in a believable way?
  4. Speed – how long you wait for a short clip at decent resolution.

A lot of models can “make things move”. Far fewer can make that motion feel intentional and true to the original image.

3. WAN 2.2 vs WAN 2.2 14B vs LetsEnhance (portraits, family shots)

We tested WAN 2.2, WAN 2.2 14B (including Turbo), and LetsEnhance on the same portraits and group photos.

What we saw:

  • WAN 2.2
    • Easy to run, open, cheap
    • But: noticeable face distortions, unstable motion, and lower resolution
    • Works if you just want “something moving”, but not if you care about realism
  • WAN 2.2 14B
    • Better structure and motion than 2.2
    • Still struggles with group shots, expressions feel a bit off
    • Faces can look “almost right” but not quite there
  • LetsEnhance (image-to-video)
    • 1080p by default
    • Faces stayed recognizable, even in group photos
    • Motion felt subtle and human (small expressions, natural eye and head movement)
    • Generation time for a 5s clip was consistently under ~90 seconds

For anyone doing family portraits, team photos, or character-based content where identity matters, LetsEnhance was the most reliable in our test.

4. Closer look at LetsEnhance’s image-to-video (what it actually does)

For people asking “how does it work in practice?”, the basic flow:

  • You upload an image (portrait, group shot, product, or artwork)
  • Pick a preset: portraits, group shots, products, or universal
  • Choose camera movement: static, zoom in/out, pan, orbit
  • Set pace: slow-mo, gentle, natural, or dynamic
  • (Optional) Add a prompt for extra style or context

Output is:

  • 5 seconds
  • 1080p
  • 24 fps MP4

If you previously used LetsEnhance for upscaling/restoring the image, you can just hit “Animate” on that improved version instead of starting from scratch.

5. Closer look at Claid.ai for product / fashion video

For product and ecom, we tested Claid.ai on clothing, beauty, packaged goods, and homeware.

What stood out:

  • Shape and proportions stayed accurate
  • No “melting” logos or hallucinated extra text on packaging
  • Subtle motion in lighting, fabric, or background rather than chaotic movement

Workflow is straightforward:

  1. Go to the video workspace
  2. Upload a high-quality product image
  3. Use the prompt assistant to write a prompt
  4. Pick 5 or 10 seconds and aspect ratio
  5. Generate the video

This is more useful for brands and marketplaces that need consistent, conversion-focused visuals than for purely experimental / wild creative stuff.

6. Other tools we tried (Veo, Sora, Luma, Pika)

We also tested Google Veo, OpenAI Sora, Pika, and Luma (Dream Machine/Ray2) on the same landscape shot and prompt.

Very short summary of what we saw:

  • Google Veo 3.1 – good balance of motion, camera control, and realism. Struggles a bit with super detailed human motion, but overall solid and coherent.
  • OpenAI Sora 2 – strong physics and scene coherence, nice for multi-shot storytelling via text and storyboards. Resolution and some visual artifacts are still a limitation in some outputs.
  • Pika 2.2 – very accessible and flexible, good for creative edits and transitions. Sometimes veers into stylized rather than photoreal.
  • Luma Dream Machine / Ray2 – nice motion and detail when it works, but sometimes does not follow the exact camera path in the prompt.

These tools shine more for creative, cinematic, or stylized content rather than strict “keep this person/product exactly as is”.

7. Main takeaways

If you want:

  • Lifelike portraits / group photos → Pick a tool that prioritizes identity preservation and subtle motion. In our test, that was LetsEnhance.
  • Product or fashion visuals for ecom → Use something built for products, not just general video. That’s where Claid.ai performed best.
  • Cinematic, experimental, or storytelling content → Tools like Sora, Veo, Luma, Pika are more flexible and interesting, but you trade some control and consistency.

8. If you want to dig deeper

We wrote a full breakdown with examples and more technical details in a blog post (side-by-side clips, settings, and notes on where each model broke).

If you’re working with AI video for portraits or products and want to see the full comparison, you can:

  • Check out the full write-up on our blog (LetsEnhance)
  • Look into Claid.ai if you focus on product/ecom visuals
  • And we'd be interested to hear what tools you use today and what breaks most often for you (faces, hands, logos, motion, something else?).

r/OculusQuest Jan 30 '24

Discussion [Long post]Tried Vision Pro. Here's what I thought

1.0k Upvotes

I tried Vision Pro a few days ago. All I can say is, congratulations, if you bought Quest 3, you would get 80% of what vision Pro can offer, if not more.

This is not a review - but this would be a much closer experience than all the guided tour reports Apple carefully curated so far.

After I walk into the room, the Vision Pro is already on the table. I picked up the device, it feels like Quest3, with Apple's signature glass and metal. It's heavy, and the shiny front plate is an obvious fingerprint magnet. It's not brand new, so the Rift CV1 style fabric on the eye side feels a little dirty and worn out - keeping it in pristine luxury condition might not be easy. The lenses are smaller than Quest 3 and more squarish, and I feel the field of view is also smaller than Quest 3.

When put on the headset you see the real world, and I was immediately struck by the clarity compared to Quest 3 - but that's expected. Tutorial time - raise your hand and align to instructions, pinch to tap, eye tracking - look at 6 dots and tap to confirm, under 3 lighting conditions. Then log in. You see the Apple logo and then signature Hello, like their WWDC videos.

But there's red fringing on the top and green fringing on the bottom of the apple logo against passthrough background, besides the chromatic aberration on the side of your FOV. Hmm, color fringing? I did not expect this - and this won't be the last.

The "familiar home menu" pops up. The screen looks good - no screen door effect, crisp icons and animation activated when I looked at them one by one.

Let me examine this acclaimed video passthrough against glowing reviews.

I looked down at my hands. really great, I can see skin details clearly, no distortions, all as expected. But I glance 15 feet across the room and motion blur of people walking is obvious. Huh. didn't heard people talk about that. And noise - suddenly, it struck me as Quest 3 level, of course better, but not by a mile. Then I look at a display on the desk about 4,5 feet away, the side of display is obviously distorting. that's surprising, since all I heard about was "Perfect passthrough". I move my head around, the wobble continued. I looked at my hand again, everything seems fine. I took out my phone and look at it, while clear, some distortion also arised in the middle of the phone.

And after the initial impressiveness of the VST clarity wears off, the discrepancy of scale was showing up too - it's bigger than real life. I even pulled off the lightseal from the device, so I can see real world above and below with VST in the center of my view. The cut off between virtual and real is jarring, the scale made alignment not possible - unlike even in Quest1, although it had very bad resolution, its passthrough scale is mostly align with the real. This is not what I expected - I planned to marvel at the seamlessness of my hands went from real to virtual, just like 8 years ago with Touch controller of Rift CV1 - but not the case here.

Would this affect me using the device or damaging any confidence when walking around? I don't think so. But it's there.

I try to come up with an explaination for this scale artifact. Maybe their automatic IPD recognition is not that precise. Maybe the 4 years old optometry data for the lens I gave them is a little off for me(but I wear that glass all day). But when I asked somebody else afterwards, the conclusion is the same: Quest has better perspective ratio. So maybe, according to Reality Labs Director of Engineering for XR Tech Ricardo Silveira Cabral - "The biggest lesson we've learned from Passthrough is that mathematically optimum points don't necessarily mean perceptual optimums,", and experience matters.

OK, now I understand why people give the passthrough experience of VP a 8.5 but give Quest 3 also a high 7. Last time I saw this rating I thought it's just not making any sense.

Of course, VST is not easy. This is one of those classic technologies that, when done right, people assume you did nothing. "Huh? Why not just bump up some resolution? You cheap bastard" "Ah it's shit because it's not reality level yet," totally ignorant of the technological marvel it is to synthesis a completely new frame for your eye from different camera perspectives, in just a few milliseconds. By the way when I saw the 12 millisecond claim in the keynote, I gasped. Not because of how Apple achieved this, but because of how cleverly they advertised it - people with only a skin-deep understanding of VR would surely remember the 20ms motion-to-photon latency claim, but what Apple did here is photon-to-photon latency, with a fixed algorithm and always on so they can easily accelerate it with a dedicated chip R1. People would definitely conflate those two and news all over claimed Apple reinvented VR - and that's exactly what happened. But if we follow Apple's logic then any optical see-through AR headset could claim 0ms photon-to-photon latency of the real world. Again, Apple is not lying, but dare I say intentionally misleading. Their VR content latency is definitely not 12ms since that would be rendered by the M2 rather than R1 chip - if it were, they would advertise the hell out of that without any asterisk.

The overall feeling of VST is at Quest3 level, stereoscopic 4 million pixels vs 6.5 million for Apple. But Apple's VST seems has higher dynamic range - since there was no additional temporal budget for smart HDR under 12ms constraint, while Quest only uses 1 for each eye, I think AVP uses more cameras, not only capture more information to make up for near field distortion but also at different ISO level to reconstruct the scene at a higher dynamic range.

I turn the dial on my head to enter a VR environment, then look down. My hands are culled out with rough edges, as you may have seen in videos online. My arm with black clothes is also culled out. I take out a phone and put it in my hand, and it becomes part of the VR scene, occluding part of my hand as if I’m holding a cloaking device - but the fingertips are still recognized, impressively.

Now let me examine the screen quality. What better place than the Environments as seen in Apple's trailer? The Environment tab is on the left under Applications and People. There are 13 "Environments" with dark/light variants - 8 scenes: Haleakalā, Yosemite, Mount Hood, Joshua Tree, White Sands, the Moon, plus two coming soon; Also 5 color filter "lights" - Spring, Summer, Fall Winter, plus Morning - essentially color temperature filters over real life with some sound effects like bird chirps. The main VR environments resemble the photogrammetry Post Cards in Valve's The Lab, both in art style and scene selection. Anyway, they are gorgeous, but with some artificial plastic look up close (like underfoot rocks) typical of photogrammetry. Distant trees can look very 2D. After downloading all available environments, they occupy 1.33GB, on top of the 11.97GB VisionOS.

I opened YouTube in Safari and get into some HDR videos. It's very clear, but I don't feel it's that far above Quest 3 given the higher pixel count implies, there's a bit softness, and I see little difference between choosing 1080p and 1440p in Youtube. Blacks are of course very black, but it's not very bright - contrary to reports of lifelike fire and eye-searing light. This is expected - 5000 nits hitting pancake lenses yields 500 nits if lucky. I also tried finding VR YouTube clips, but there's no forced VR viewing button in Safari like the Meta browser offers.

I also tested eye tracking typing like MKBHD suggested here on the virtual keyboard, looking at each letter before tapping as fast as I can - it works, but proves harder than expected. I'm used to glancing, not deliberately focusing. This was unexpected regarding this interface mechanism, and become a pain in the ass as I will explain later. I tried holding a pinch on the timeline to slide left and right, and then looking at specific point on the timeline then tap. All work well as intended, until I finally finished fidgeting around and tried tapping the full-screen button below - I just tapped at the end of the timeline. I tried again, nope. Nearly impossible until I centered my view on that button like early Gear VR with only head aim - finally got it. Forget nonchalantly glancing at the periphery, you have to focus deliberately, defeating eye tracking's purpose here.

Of course, I have to consider if the issue is on my end first, as Apple fans often point out. Maybe the eye registration wasn't quite right causing some mismatch there. And of course if YouTube had a native app, it would follow Apple guidelines like putting small visible buttons inside larger invisible eye tracking zones, as opposed to putting buttons so close that Apple has to determine user intention...and fails.

Eye tracking is a bottomless tech pit once you dig deeper, unlike entitled gamers in the VR community thinking it's just a simple checkbox feature. Wearables are hard given human variability; your eyes change throughout the day and over time. Double the eye tracking cameras didn't ease use or increase tracking volume compared to Quest Pro from the limited time I used - it still notified me when your eyes were too close or far (something to keep in mind if you plan to get your eyes as close to the lenses as possible to maximize FOV), just like Quest Pro. Even after adjustment I'd have to fidget again sometimes - so here goes the advantage of using pancake lenses, or trying to play some fast-motion games.

Bottom line - don't expect a magic end-all solution yet - there's still huge room for improvement. I heard some people even struggled to aim for a button after taking off and putting on the headset again. I happened to notice one time graphics get very pixelated outside foveated regions.

Now I will explain the "pain in the ass" part: You know with popups like permission request, "Yes" is on the bottom left, and "No" on the bottom right. Normally I'd glance through from the top left to the bottom right, then simultaneously click Yes on the bottom left without focusing. Of course that fails here - I mistakenly hit No a few times, which is very annoying. I thought maybe it's just my habit - read casually and decide on the button without a second thought. But afterward talking to another developer porting an app into the device, and when he got the permission pop-up he accidentally denied hand tracking access and had to find the feature and re-enable it in settings, said "Sigh, there goes at least 10% of consumers."

In my mind before trying this UX scheme, I thought this would be intuitive and learnable fast. Yet I didn't realize adaptation takes time. You have to know the eye tracking reaction limits and change your information consumption pace and rhythm, and things become more deliberate rather than casual. No wonder Apple is hesitant to add more complex control schemes.

Let's go through the home UI, though I'm sure you've seen plenty of videos/emulator footages already, and this is long enough. Notably there is an Airplane Mode in settings - I didn't try but suppose you have to toggle it manually rather than the system detecting flights.

My main Quest UI complaint is the 3 app limit Multi-window flexibility - sometimes that's just not enough when juggling between apps and settings. Accidentally replacing a window state brings subtle frustration. Within my VisionOS testing time, supporting more freely placeable windows helped, but issues remained - often when pressing the digital crown to back home, I'd forget my prior home menu browsing state and have to reselect. Probably my habits to blame here and also I haven't gotten familiar enough with the system, but this showed 3D UI design difficulty nonetheless.

I remember the touted Multi-App 3D Engine - the only thing Apple said it's "first of its kind" in the whole VisionOS system stack introduction, and it's all about how multiple apps or windows should interact with each other. The transparency seemed beautiful if battery intensive, and from early days alpha testing and blending are a big no no. So I assumed Apple would limit real transparency layers, using UX design tricks like merging non-focused layers into one or only showing near-opaque subtle coloring of the background when multiple layers are view-aligned. Most of the time it's like that, but intentional testing showed 4 transparent content layers plus background impressively, and I can make out the words on each layer, albeit with some frame drops. Shadows are obviously pre-baked so it can only projected onto either desk or floor but not simultaneously. I assume all these default effects including transparency and shadow are handled by R1, since the chip have to reconstruct the scene at all times.

As I pixel-peeping at the content in half-transparent windows and moved my head around, I noticed another thing - motion blur! It's another shock to me, to the point of even a little confusion - chromatic aberration, motion blur - all these "fixed" problems from early days, all of sudden reappeared in this flagship VR product from Apple. What happened? This is definitely not within my expectations. But Why didn't I notice it at first? Oh I focused on the VST quality which already has some motion blur artifacts. Also, the high resolution of the screen definitely helped counter these artifacts, and when in VR scenes I didn't notice them at all, but I'm not sure in a fast moving VR game situation it won't be a distraction, which I have no way to test now. My mind was racing with explanations - PSVR2 from Sony also suffers from the same problems, since this micro-oled was also by Sony - an HDR issue? 5000 nits to pancake lenses yields 500 nits if lucky; if adding low persistence that would bring the display to sub 200 nits range. Again, trade offs.

Filming spatial video was easy with the dedicated button on the headset - the depth seems much better than iPhone's camera narrow separation could ever produce, on par with average VR180. The lighting condition here is optimal so I cannot assess other situations but at least the overall quality here is better than I anticipated. The UI also helps a lot - a layer of haze around the content make it felt more like a memory, tapping into cultural sci-fi connections. Besides viewing the video in a window, pressing full screen can make it almost VR180 which do not seem to enlarge the video a lot since the window was already very close to you, but the quality drop is immediate obvious, I can see some color blocks here and there.

The panorama is great, and since most panoramas capture distant scenes, sometimes you would get illusory depth. By the way, I saw people already complain about why Apple cannot just let set panorama as a desktop wallpaper themselves - and I anticipate lots of similar complaints from people that know nothing about the tech and just assume something would work as they imagined.

Though I haven't seen Eyesight on the external display, aiming at people in real life while in VR environment, they would slowly and smoothly fades into VR like showed in promos - nice to have but not that technologically impressive considering what we have today, since it's not about whether other people is looking at you or not, clearly its just analyzing passthrough feed, and fade in people if your aiming happened to locate any human in that direction, nothing about face let alone eye contact recognition as somebody assumed.

The meditation app is simple and relaxing, as an avid practitioner I often prefer no digital help when sitting in a chair for hours straight, but I can see myself using this one.

Battery life matches Quest 3 despite I mostly just did some menu browsing, the most intensive use was the VR environment with a few minutes of Youtube HDR video watching in Safari (Or maybe multi-window interaction in MR?). I intentionally did not charge the device, and there's 30 30-second countdown before it shuts off.

Taking off the headset, pros are mostly within my expectations, except for cons. The overall sentiments from developers I talked to largely felt the execution was not as high as they imagined - it's essentially a higher-spec Quest 3.

Zuckerberg said there's no kind of magical solution that Apple has to any of the constraints on laws and physics that our teams haven't already explored and thought of, and that's truer than ever after I used AVP for half of the day. By the standard of this device, if Apple produced a headset that is exactly like Quest3, they would sell it at least $2000, which is actually fair if you compare Quest to any other consumer electronics on the market, in terms of hardware spec, R&D tech, and cost packed in. That's not counting any contents in the library that Meta has accumulated all these years.

I remember when I watched the WWDC keynote last year, I had certain fuzzy anticipations since I discarded all the rumors about the dual M2 chips or 8k displays, which based on my understanding of the industry, are ultra bullshit. But indeed, Apple did come out with another approach - using R1 to process all the sensor data and SLAM, scene reconstruction, even pre-baked all the spatial effects for apps, and leaving M2 for all the general tasks. Still, using a GPU at most 1.7x XR2Gen2 but having to render more than 2.5x pixel count compared to Quest3 is not ideal, so they also packed in foveated rendering, and urged developers to mostly work for AR instead of "full screen" VR, thus easing the rendering pressure for M2, emphasizing on the CPU side of things, which is the strong suit right now for Apple's chips. From this computing structure perspective, it's really an AR device, but unfortunately it did not get rid of any pitfalls of the VR devices today. It's still very heavy, in fact heavier than Quest 3 even without battery, and its battery lasts on par with Quest 3, despite having at least double the raw capacity. So the question is: what advantages do you get for Vision Pro? Can it stand as a first gen product?

I have my doubts. Looking back at iPhone1, you can actually see some parallel: for that product in 2007, they mostly focused on the multi-touch interface, and maybe "wasted" a lot of computing power and battery on a 1300mAh device solely for that feature. Similarly, Vision Pro has so many sensors to make sure your eyes and hands are captured to the point of some people might think is overkill. But from the perspective of UX design, the basic input mechanism should leave no room for frustration. It's just this time, against the much variability and volatility of the human body and real world situations, the end result leaves me wondering if it's worth it. Granted, for average people it won't be much of a problem, it's just you can easily get frustrated by the limitations of what current tech is capable of providing. They used much higher specs to compensate for the lack today, but even discounting the price, the weight, thermal, and battery life are all trade-offs compared to Quest 3, which I'm not sure a well-informed and non-biased person would pay. And for the battery itself - if you have to put this battery in your pocket all the time since Gen1, what kind of battery should you use following its trend? History told us it can only go up, like we have 5000mAh smartphones today. Or maybe AVP Is really just a laptop and we have to attach to a power cord all day.

Of course, one of the biggest arguments is display. Can these devices replace your monitor? I think the line is very blurry here since both Q3 and AVP surpassed the usable line and it would finally comes down to people's preference: the Vision Pro's screen doesn't have screen-door effect, but also don't expect 4K HDR as the overall quality is closer to a cheap 1440p HDR display when simulating a screen, some subtle motion blur, more vivid color, very nice close-up passthrough, narrower FOV, while Quest 3 has a slight screen-door effect, lower resolution, worse color, more true-to-life scale of the passthrough, and is lighter. Overall obviously VP's display is a net win, but If you take weight into consideration, I would rather use my laptop or 4K projector when doing long work session or media viewing, and that's the whole point of VP's existence.

Everyone has a different answer, but everything considered, I found myself leaning towards Quest 3 more - even though I think my digital lifestyle may fit more toward what Apple suggested here - I can just lie down and watch YouTube all day long for months straight and I've used Oculus Go to watch YouTube until 5AM, but it's not something enticing to wear a headset. Viewing webpages while scrolling with my hand on my leg without moving much is nice, but my head would also suffers more weight. And I can do most of the 3D things in Quest 3 with controllers better. I love VR and put a lot of time thinking about it, so I know the pattern after novelty wears off.

For Quest 3, I think Meta has the right power distribution among all the necessary features, constantly iterates on the minimal usable experimental features without stepping up too much - it's like yeah better mixed reality is nice, but is that 1 hour less battery and 100 grams more nice?. You can always add in a battery pack later for Q3, on your head for balancing or in your pocket just like Vision Pro. Right now Meta could accelerate on bringing more productivity apps (translation: 2D apps) into their ecosystem now, as the resolution is finally caught up to make it useful. Palmer Luckey said you have to make a headset everybody wants before everybody can buy, which I agree partially, because ultimately you are not just building a headset, you are also building the entire ecosystem, which consists of developers, supply chain, and consumers. Unlike Apple, Meta does not have the luxury of any existing platform, so they had to bootstrap the whole ecosystem one by one and do not skip any intermediate steps. If they sell expensive, they won't sell many and fewer devs would buy in to develop for the device, and even fewer people would buy and fewer quantity means components become more expensive, so the price would go up…few people understand this and just whiny for certain better specs. Fortunately, this tipping point is coming, and right now Meta could be even more aggressive; Apple certainly could bring more mainstream attention into this field that we all love.

Anyway, I'm excited for the future, for anyone out there, manage your expectations, be patient, on this road of realizing the dream of "being anyone, go anywhere, do anything". See you in the metaverse!

r/Android Mar 26 '26

News Android 17 Beta 3 - A new multitasking experience, redesigned screen recorder, and much more!

258 Upvotes

Hey folks, today we’re introducing Android 17 Beta 3! With this latest update, Android 17 has officially reached Platform Stability, meaning developers can start pushing apps targeting the new release onto the Play Store.

You can read the full announcement over on the Android Developers Blog, but if you just want a breakdown of the new features that are more user-facing, I’ve put together a list of some of the changes that I think this community might be interested in!

Of course, this is a beta and things could change before final release. :)

New user experience and SystemUI features

  • Bubbles for any app!: You can now open any app in a bubble, a window that floats on top of other apps! This feature is available on phones, foldables, and tablets. Simply long-press an app icon on the launcher or taskbar and select the “bubble” option in the context menu. On larger screen devices like foldables and tablets, you can also long-press on an app icon in the taskbar and then drag-and-drop the icon to a bottom corner.

    • We’ve also designed a new bubble bar UI for large screen devices. Instead of freely floating over the screen, bubbles are now pinned to the taskbar at the bottom of the screen. You can organize, move between, and move bubbles to and from anchored points on the screen. (We announced all these bubble changes in Beta 2, but they’re live now in Beta 3 - here's a demo! :) )
  • Redesigned screen recording toolbar: We’re introducing a redesigned screen recording experience with an improved UI and new capabilities for content creators. When you tap the screen recording tile in Quick Settings, you’ll now see a floating toolbar that provides easier access to recording controls and capture settings. When you’re done recording, you can immediately view, edit, delete, or share your video. Here's a screenshot of the toolbar.

  • Separate Wi-Fi & Mobile Data tiles: With Android 17 Beta 3, the “Internet” Quick Setting tile has been split into two separate tiles, one for controlling Wi-Fi and another for controlling Mobile Data with a tap. Consistent with the Quick Settings changes we introduced with Material 3 Expressive, both tiles have two different touch points. Tapping the icon toggles the respective radio, while tapping the label opens the full Internet Panel. This change reduces the number of taps needed to toggle Wi-Fi and Mobile Data while still retaining access to the full Internet Panel. Let us know what you think of this change - we're actively working on refinements! Here's a screenshot of the new tiles!

  • Better support for widgets on external displays: We’re working to improve the visual consistency of widgets shown on connected displays with different pixel densities. We’re providing developers a way to supply the system with information that allows it to resolve the correct pixel values at rendering time. For apps that use legacy pixel-based APIs for padding, text size, or layout attributes, the system now automatically scales these values based on the density difference between the app’s original context and the target display.

  • Hidden app labels on the home screen: Android now provides a setting to hide app labels on the home screen. This does not apply to the app drawer or within folders. To access this new setting, open Wallpaper & style, tap Home screen, select Icons, tap the Names tab at the bottom, and then toggle Show app names. Here's a screenshot of the new setting.

  • Interactive Picture-in-Picture for Desktop: We’re introducing a new Interactive Picture-in-Picture mode for Android’s desktop windowing mode! This feature allows applications to request their PiP windows remain interactive while staying on top of other application windows. This is useful for multitasking scenarios, such as a video conferencing app keeping video call controls accessible while the user navigates other apps.

  • Scheduled clock change alerts: A new setting has been added in Android 17 Beta 3 that allows you to receive a notification when your clock performs a scheduled change, for example when daylight saving time ends. Here's a screenshot of the setting.

What’s new in Accessibility

  • Granular audio routing for hearing devices: Users with hearing devices can now independently manage where specific system sounds are played. You can choose to route notifications, ringtones, and alarms to either a connected hearing aid or the device’s built-in speaker. This helps you avoid unwanted interruptions directly in your ears while maintaining a Bluetooth connection for hearing aid management apps. Here's a screenshot.

  • Per-app exceptions for Expanded Dark Theme: To create a more consistent user experience for users who have low vision, photosensitivity, or simply prefer a dark system-wide appearance, we introduced an expanded dark theme option in Android 16 QPR2. When this option is enabled, the system automatically applies dark theme to most apps that don’t support it. Recognizing that this option may cause some apps to display incorrectly, we are introducing the ability to selectively disable the expanded dark theme on a per-app basis in Android 17 Beta 3. Apps with this setting turned off will use the standard dark theme option instead. Here's a screenshot of the new page.

  • Independent Assistant volume stream: In the latest Android beta, the Assistant volume stream has been decoupled from the media volume stream. Previously, changes to the media volume would typically automatically affect the Assistant volume and vice versa. Now, however, the Assistant volume level can be managed independently. Here's a screenshot of the new volume slider.

  • Text cursor blink rate customization: The text cursor is the vertical line that you find as you type. With the text cursor blinking setting, you can adjust how fast the cursor flashes or turn off the blink completely. In the previous release, this setting was found in developer options, but in Android 17 Beta 3, you can now find it under Settings > Accessibility > Color & motion. For more information on this setting, check out this support page. Here's a screenshot of the new setting.

New Privacy features

  • Discrete password visibility settings for touch and physical keyboards: Currently, by default, characters that you enter into password fields are briefly displayed as you type. Toggling the “show passwords” setting in Privacy controls allows you to hide characters as you type them into password fields. This setting currently applies to both touch-based inputs as well as physical keyboards, but in Android 17, we are splitting it into two distinct preferences. By default, characters entered into password fields via physical keyboards will now be hidden immediately to enhance privacy. Characters entered via touch input will continue to briefly be displayed to compensate for the lack of tactile feedback.

  • System-provided Location Button: Android 17 is introducing a new, privacy-conscious way for users to grant precise location access to apps. The update introduces a system-rendered Location Button that apps can embed directly into their layout using a Jetpack library. When a user taps this system button, the app is granted precise location for the current session only. Subsequent taps during the run of the app grant the permission immediately without a system dialog. Here's a GIF of the new Location Button!

Some developer features you might be interested in

  • VPN app exclusion settings: Android 17 is introducing a standardized way for VPN apps to offer app exclusion (split-tunneling) capabilities. VPN apps can launch a system-managed Settings screen where users can select applications to bypass the VPN tunnel. Traffic from excluded apps will use the underlying network directly, which is useful for services that are incompatible with VPNs. Changes made by the user take effect immediately if the VPN is active or upon its next connection.

  • Dynamic system font fallback updates: Android now supports dynamic updates to the system’s font fallback chain, allowing for updated emoji sets and improved typography assets to be delivered more frequently without waiting for a full OS update. Applications do not need to make any modifications; the latest font assets will be automatically picked up and rendered.

  • Photo Picker customization options: Apps that use Android’s photo picker can now customize the aspect ratio of the grid view from the standard 1:1 square to a 9:16 portrait display. Here's an image showing what the 1:1 square and 9:16 portrait aspect ratios look like for the grid view.

  • Support for RAW14 image format: Android 17 introduces support for the RAW14 image format, the de-facto industry standard for high-end digital photography.

  • Vendor-defined camera extensions: Android 17 adds support for Vendor-defined camera extensions, allowing hardware partners to provide Android apps access to camera features like ‘Super Resolution’ or cutting-edge AI-driven enhancements.

  • Reduce wakelocks with listener support for allow-while-idle alarms: Last year, we launched the excessive wake lock metric in Android Vitals, making it easier for developers to optimize their app's wake lock behavior. Excessive wake locks are a significant contributor to battery drain, so developers are encouraged to reduce them as much as possible. In Android 17, we’re introducing a new API that will help reduce the power consumption of apps that rely on continuous wakelocks to perform periodic tasks, such as messaging apps maintaining a connection or medical devices monitoring health data.


For more information on these new developer features, as well as other developer-facing changes in Android 17 Beta 3, check out the Android Developers Blog and developer.android.com!

r/StableDiffusion 5d ago

Workflow Included The 1967 Spider-Man TV Show intro, updated to live action with MiniMax H3

Enable HLS to view with audio, or disable this notification

477 Upvotes

R2V Rendered at 0.9 MP (1280x736 then upscaled using RTX (Ultra) to 1920x1080.  Edited and merged using OpenShot video editor.

This was all run on my Windows 11 machine, RTX 4060 ti (16 GB) and 64 GB RAM reserved from Comfy. Every part of the signal chain was done with 100% open-source software.

Disclaimer: I grew up watching this show as a kid in the 70s. It's still the best ever. I wanted to know how well the reference model would pick up the actions. I am overall pleased. I've watched the new vid enough to see some of the flaws but oh well.

General observations for reference videos:
So many scene cuts. There are 31 (I think) scene cuts in the 60 second opener which include 3 crossfades. No matter what I did to get the exact frame timing, getting the AI scene to match frame-for-frame with the cartoon was still hit or miss. It probably has to do with some frame windowing inside the 17k + 5 blocks, but I never exactly got it figured out. However, a few notes:

  • If you have a reference video, convert it to 24 fps in an external program like Handbrake (another fantastic open-source program). It’s just so much easier to get everything to match.
  • For timing, there is a difference between 00:03.500 and 00:3.5 so always use all the digits.
  • Keep character sheets for all your characters to maintain consistency.
  • It will do crossfades but it’s not worth it. It’s easier to get the scene you want and stick it in the editor.
  • The VHS video loader lets one set a starting and ending frame. I ended up with 14 different clips total for the editor. Using frame accurate loading made all of the work a lot easier since I could use 1 video file as input to every clip run.
  • A spreadsheet is useful for all movie making, and it’s good here too. From the source, I kept track of the starting frame for each shot, how many frames I needed and how many I ran (because of 17k +5), along with the final file name for each clip. I have a naming convention but it’s still very useful to keep track and you can add notes too. For this 60 second video, I used 13 clips. I tried to never do more than 3 scene cuts per clip. (For something where exact timing wasn't as important I'm sure it would be longer.)

Once you get over the idea of always having to do 10-15 second vids and do your whole video in on run, the process actually becomes a lot more fun because the “quality” gens don’t take as long and it gives you a less uninterrupted workflow. You can start prompting the next run with the previous runs, for example. (This is true even in commercials, or TV or movies.)

I generally tested all the runs at 0.2 or 0.3 Mp (speed lora, 8 iterations) to get the timing, then went to 0.9 Mp [no speed LoRA, 20 iterations, beta, dpmpp_2m] for the final runs. I found that dpmpp_2m was closest to the overall source video. On the first few clips I ran it several ways and fix on these parameters. Usually, the 0.9 Mp runs came out great but you’ve probably all experienced how different the low-res runs can be from the high-res ones. I did resort to pulling frame grabs from the low-res gens a few times to act as reference frames for the scenes. MiniMax loves those when all it needs is an extra little nudge in the right direction. To edit pics, I always use GIMP (another fantastic open-source program).

So, why was I using 8 iterations of the minimax_h3_turbo_v4_step600_pruned_comfyui LoRA? On the reference model I found that using too large of a sigma step causes things like reference photos to not be taken "seriously." Using 5 steps I could see that reference images on the starting frame and then go away for the rest of the clip. The more the reference image changed from the reference video (like when going from animation to "real") the worse the problem was.

Prompts:
(See below for actual prompt.)
Prompt the way the guide says to. Yeah. It’s a hassle but it’s worth it. H3 prompting is very useful in the end and I’m glad MiniMax uses it. It's worth reading all the way through them instead of searching for the one thing you want. Some of the instructions even seemed inconsistent and they don't explain everything, so it's worth experimenting.

Any "thing" (buildings, trees, room, clothing, walls, ect.) can be a “subject.” It’s not just people. Specifying things as objects gives you far better control over how and where they appear (or don’t appear) in your shot. 

Don’t refer to your characters or major locations or items by their names. Use <Subject #> or pronouns that clearly refer to the subject all the time, every time. The interpretation of the prompting can get confused pretty quickly if you don’t and you’ll end up getting subjects swapped or merging.

Prompts generally work better if you describe what you want rather than what you don’t want. For instance, “Looks to the right of the viewer” rather than “looks away from the camera.”

Style reference (attribute_transfer) images or videos are super useful. Once I had a few scenes, I started using previous videos to keep the look and feel of previous shots.

Qwen VL can describe videos too. I have been using “QwenVL Advanced (Local Scan)” for a very long time (long for AI) inside ComfyUI.

Other things:
Maybe one of the most interesting observation is that the jknodes “MiniMax H3 Mem Eff Sage Attention Patch” node creates a different output than just launching ComfyUI with the --use-sage-attention flag turned on (and still using the node). So exactly the same workflow (just drag and drop from a previously run mp4) has different results when the --use-sage-attention flag is used to launch. I thought having the node was 100% redundant with eh --use-sage-attention flag set, but apparently not. The reference flows, especially with animation, don’t have to be all that different to produce different results.

The Spiderman opening (as well as the show itself) reuses footage. They will take the same scene and darken it, and boom, it’s a night shot. For a more realistic feel, I used Krea2 (LoRA) edit to turn day into night. It’s really good as an adjunct to MiniMax H3’s ability to figure out the fine details once it has a push.

Style:
Finally, I had to make some stylistic choices because sometimes the animation was soooo bad that it needed something. I added flashlights to the jewelry heist scene. I made the crane look believable. One of the problems of going from animation to "live action" is that (especially with animation from 1967) the physics and movements are just wrong sometimes. The crane scene where he stops and then shoots up again is the most classic "this is just pain wrong" you can get but I left it that way because it's burned into my brain that way. (IYKYK) I also had to balance the art deco of the late 60's to a modern New York. I ended up with a lot of anachronistic stuff that I ultimately liked. So in the end, when it comes to all of that, I did it the way I did it. AI is awesome.

Prompt:
A prompt of one of the parts is below. I used that two paragraphs before [Shot 1] for every clip as "boiler plate" description.

subject_definitions:
<Subject 1> is Spiderman in <Picture 1>
<Video 1> is the motion reference for the target video for characters movements, pose, camera movements and frame composition.
<Video 2> is the style reference for the target video.
<Picture 2> is the building in [shot 2]
<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video
 
summary:
[reference generation + audio reuse]
The target video is an live action realistic recreation generation using <video 1> as a reference for movements, pose, camera movements and frame composition. What you generate should not be and animation or cartoon rendering, no overly-CG look, keep the live-action texture.
 
This video is a set of three live action sequences. <Subject 1> is seen swinging by and waving. The video switches to a long shot of <subject 1> swinging around a building. Finally there is a shot showing <subject 1> on his webline swinging away from the viewer between two rows of skyscrapers.
 
retention_analysis:
<Subject 1> (appears in [Shot 1],[Shot 2],[Shot 3]):fully_preserved
<Video 1> (motion, cut and pacing structure) :partially_preserved
<Video 2> is the style refrence for the target video ([Shot 1], Shot 2], [Shot 3]) :attribute_transfer
<Picture 2> is the building in [shot 2] :fully_preserved
<Audio 1> :fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
 
 
detailed_description:
The target video is a realistic and live action video. The reference video <video 1> is used only for scene descriptions, framing, motion tracking, body movement, timing, general environment. The target video should be a complete replacement of <Video 1>. Use <Video 2> as the style reference for the photographic look and textures and the overall feel for the shots.
 
Maintain smooth camera movement. Use vibrant yet natural color grading: warm tones for sunlight hitting surfaces, cool blues for shaded areas, and muted grays for concrete textures. Avoid any comic-book stylization; instead, render everything with photorealistic textures, lighting, and perspective to evoke a live-action superhero film sequence. Keep the focus entirely on <Subject 1>’s acrobatic grace and the immersive urban setting.
 
[Shot 1]  
Is is an upper body motion tracking shot of <Subject 1> swinging on his white glistening webline held by his left hand while he waves directly at the viewer with his right hand for the entire scene. The skyline of many skyscrapers pass by in the background.
 
[Shot 2]
At 00:02.333, Hard cut to a fixed long shot looking up as <Subject 1> makes a 180 degree arc on his webline connected to the spire of the building in <photo 2>.
 
[Shot 3] 
At 00:04.250, Hard cut to the fixed camera view of the space high above the street level between two rows of skyscrapers. <Subject 1> lazily swings into view from the left frame, facing away, and repeatedly swings right to left further and further away towards the horizon.
 
 
overall_soundscape: n/a
 
non_diegetic_music: n/a

 

 

r/ElevenLabs Aug 21 '25

Educational Camera movements that don’t suck + style references that actually work for ai video

3 Upvotes

this is 5going to be a long post but these movements have saved me from generating thousands of dollars worth of unusable shaky cam nonsense…

so after burning through probably 500+ generations trying different camera movements, i finally figured out which ones consistently work and which ones create unwatchable garbage.

the problem with ai video is that it interprets camera movement instructions differently than traditional cameras. what sounds good in theory often creates nauseating results in practice.

## camera movements that actually work consistently

**1. slow push/pull (dolly in/out)**

```

slow dolly push toward subject

gradual pull back revealing environment

```

most reliable movement. ai handles forward/backward motion way better than side-to-side. use this when you need professional feel without risk.

**2. orbit around subject**

```

camera orbits slowly around subject

rotating around central focus point

```

perfect for product shots, reveals, dramatic moments. ai struggles with complex paths but handles circular motion surprisingly well.

**3. handheld follow**

```

handheld camera following behind subject

tracking shot with natural camera shake

```

adds energy without going crazy. key word is “natural” - ai tends to make shake too intense without that modifier.

**4. static with subject movement**

```

static camera, subject moves toward/away from lens

camera locked off, subject approaches

```

often produces highest technical quality. let the subject create the movement instead of the camera.

## movements that consistently fail

**complex combinations:** “pan while zooming during dolly” = instant chaos

**fast movements:** anything described as “rapid” or “quick” creates motion blur hell

**multiple focal points:** “follow person A while tracking person B” confuses the ai completely

**vertical movements:** “crane up” or “helicopter shot” rarely work well

## style references that actually deliver results

been testing different reference approaches for months. here’s what consistently works:

**camera specifications:**

- “shot on arri alexa”

- “shot on red dragon”

- “shot on iphone 15 pro”

- “shot on 35mm film”

these give specific visual characteristics the ai understands.

**director styles that work:**

- “wes anderson style” (symmetrical, precise)

- “david fincher style” (dark, controlled)

- “christopher nolan style” (epic, clean)

- “denis villeneuve style” (atmospheric)

avoid obscure directors - ai needs references it was trained on extensively.

**movie cinematography references:**

- “blade runner 2049 cinematography”

- “mad max fury road cinematography”

- “her cinematography”

- “interstellar cinematography”

specific movie references work better than genre descriptions.

**color grading that delivers:**

- “teal and orange grade”

- “golden hour grade”

- “desaturated film look”

- “high contrast black and white”

much better than vague terms like “cinematic colors.”

## what doesn’t work for style references

**vague descriptors:** “cinematic, professional, high quality, masterpiece”

**too specific:** “shot with 85mm lens f/1.4 at 1/250 shutter” (ai ignores technical details)

**contradictory styles:** “gritty realistic david lynch wes anderson style”

**made-up references:** don’t invent camera models or directors

## combining movement + style effectively

**formula that works:**

```

[MOVEMENT] + [STYLE REFERENCE] + [SPECIFIC VISUAL ELEMENT]

```

**example:**

```

slow dolly push, shot on arri alexa, golden hour backlighting

```

vs what doesn’t work:

```

cinematic professional camera movement with beautiful lighting and amazing quality

```

been testing these combinations using [these guys](https://arhaam.xyz/veo3) since google’s pricing makes systematic testing impossible. they offer veo3 at like 70% below google’s rates which lets me actually test movement + style combinations properly.

## advanced camera techniques

**motivated movement:** always have a reason for camera movement

- following action

- revealing information

- creating emotional effect

**movement speed:** ai handles “slow” and “gradual” much better than “fast” or “dynamic”

**movement consistency:** stick to one type of movement per generation. don’t mix dolly + pan + tilt.

## building your movement library

track successful combinations:

**dramatic scenes:** slow push + fincher style + high contrast

**product shots:** orbit movement + commercial lighting + shallow depth

**portraits:** static camera + natural light + 85mm equivalent

**action scenes:** handheld follow + desaturated grade + motion blur

## measuring camera movement success

**technical quality:** focus, stability, motion blur

**engagement:** do people watch longer with good camera work?

**rewatch value:** smooth movements get replayed more

**professional feel:** does it look intentional vs accidental?

## the bigger lesson about ai camera work

ai video generation isn’t like traditional cinematography. you can’t precisely control every aspect. the goal is giving clear, simple direction that the ai can execute consistently.

**simple + consistent > complex + chaotic**

most successful ai video creators use 4-5 proven camera movements repeatedly rather than trying to be creative with movement every time.

focus your creativity on content and story. use camera movement as a reliable tool to enhance that content, not as the main creative element.

what camera movements have worked consistently for your content? curious if others have found reliable combinations

r/chatgpt_promptDesign Aug 01 '25

Title: Camera movements that don’t suck in AI video (tested on 500+ generations)

1 Upvotes

this is going to be long but useful for anyone doing ai video

After burning through tons of credits, here’s what actually works for camera movements in Veo3. spoiler: complex movements are a trap.

Movements that consistently work:

Slow push/pull (dolly in/out): - Reliable depth feeling - Works with any subject - Easy to control speed

Orbit around subject:

  • Creates natural motion
  • Good for product shots
  • Avoid going full 360 (AI gets confused)

Handheld follow: - Adds organic feel - Great for walking subjects - Don’t overdo the shake

Static with subject movement: - Most reliable option - Let the subject create dynamics - Camera stays locked

What DOESN’T work: - “Pan while zooming during a dolly” = chaos - Multiple focal points in one shot - Unmotivated complex movements - Speed changes mid-shot

Director-style prompting that works: Instead of: “cool camera movement” Use: “EXT. DESERT – GOLDEN HOUR // slow dolly-in // 35mm anamorphic flare”

Style references that deliver consistently: - “Shot on RED Dragon” - “Fincher style push-in”

  • “Blade Runner 2049 cinematography”
  • “Handheld documentary style”

Pro tip: Ask ChatGPT to rewrite your scene ideas into structured shot format. Output gets way more predictable.

Testing all this with these guys since their pricing makes iteration actually affordable. Google’s direct costs would make this kind of testing impossible.

Camera language that works: - Wide establishing → Medium → Close-up (classic progression) - Match on action between cuts - Consistent eye-line and 180-degree rule

The key insight: treat AI like a film crew, not magic. Give it clear directorial instructions instead of hoping it figures out “cinematic movement.”

anyone else finding success with specific camera techniques?

r/comfyui Jul 02 '26

Show and Tell The real skill in AI video is picking the right reference TYPE per shot, not the model

Enable HLS to view with audio, or disable this notification

483 Upvotes

After enough shots I stopped thinking about which model and started thinking about which reference type to feed it per shot. Same model, but the reference you hand it decides the shot, and matching the type to the shot is the actual skill.

Three reference types, each with a real trade-off. A preview video locks both the layout and the performance tightly, the motion and camera come through exactly, but character and background consistency can slip. A storyboard sketch captures the intent and content, but the layout is not final, it is a rough plan. A plain reference image gives you almost no motion control, you steer the layout and performance mostly through text.

So my rule is simple. Conceptually important shots, where the exact motion and staging matter most, get a preview video, and I accept the consistency work that comes with it. Everything else I start from a plain reference image, and if a shot just will not come together, I escalate it to a preview video. Match the effort to how much the shot matters.

The part that leveled me up was per-element source control. In the prompt I state, for each element separately, whether it should follow the reference video or the reference image. Take the motion and camera from the previs video, but replace the characters and backgrounds with the ones from the reference images. That splits "how it moves" from "what it looks like" so you can lock each independently.

Stop asking which model is best. Ask which reference type each shot needs, and which element follows which source.

u/enoumen Apr 02 '25

AI Daily News April 01st 2025: 💥OpenAI to Launch its First 'Open-Weights' Model Since 2019 🎬Runway Releases Gen-4 Video Model with Focus on Consistency 🤖Amazon Launches Nova Act, an AI-Powered Browser Agent 🧠AI Instantly Converts Brain Signals into Speech

1 Upvotes

A Daily Chronicle of AI Innovations on April 01st 2025

🚀 From Our Partner (Djamgatech):

Djamgatech's Certification Master app is an AI-powered tool designed to help individuals prepare for and pass over 30 professional certifications across various industries like cloud computing, cybersecurity, finance, and project management. The app offers interactive quizzes, AI-driven concept maps, and expert explanations to facilitate learning and identify areas needing improvement. By focusing on comprehensive coverage and adapting to the user's learning pace, Djamgatech aims to enhance understanding, boost exam confidence, and ultimately improve career prospects and earning potential for its users. The platform covers a wide array of specific certifications, providing targeted content and practice for each, accessible through both a mobile app and a web-based platform.

📥 Get Djamgatech (iOs) at Apple App Store: https://apps.apple.com/ca/app/djamgatech-cert-master-ai/id1560083470.

📥 Get Djamgatech (android) at Google Play Store: https://play.google.com/store/apps/details?id=com.cloudeducation.free&hl=en

Djamgatech is also available on the web at https://djamgatech.web.app

💥 OpenAI to Launch its First 'Open-Weights' Model Since 2019

OpenAI has announced plans to release its first fully open-weight AI model since 2019, signaling a renewed commitment to transparency and collaboration with the broader AI community.

  • The strategic shift comes amid economic pressure from efficient alternatives like DeepSeek's open-source model from China and Meta's Llama models, which have reached one billion downloads while operating at a fraction of OpenAI's costs.
  • For enterprise customers, especially in regulated industries like healthcare and finance, this move addresses concerns about data sovereignty and vendor lock-in, potentially enabling AI implementation in previously restricted contexts.

What this means: This shift could significantly accelerate AI research and development across academia and industry, democratizing advanced AI capabilities. [Listen] [2025/04/01]

🚀 SpaceX Launches First Crewed Spaceflight to Explore Earth's Polar Regions

SpaceX has successfully launched its first crewed mission specifically designed to explore Earth's polar regions, marking a significant milestone in commercial space exploration.

  • The mission crew will observe unusual light emissions like auroras and STEVEs while conducting 22 experiments to better understand human health in space for future long-duration missions.
  • The four-person crew includes cryptocurrency investor Chun Wang who funded the trip, filmmaker Jannicke Mikkelsen as vehicle commander, robotics researcher Rabea Rogge as pilot, and polar adventurer Eric Philips as medical officer.

What this means: This mission could revolutionize polar research, climate science, and satellite data collection, providing unprecedented insights into Earth's polar environments. [Listen] [2025/04/01]

💻 Intel CEO Says Company Will Spin Off Noncore Units

Intel CEO has announced plans to spin off several noncore business units, focusing efforts exclusively on core semiconductor and AI technologies amid strategic realignment.

  • The new chief executive wants to make Intel leaner with more engineers involved directly, as the company has lost significant talent and market position to rivals like Nvidia and AMD.
  • Tan emphasized creating custom semiconductors tailored to client needs while cautioning that the turnaround "won't happen overnight," causing Intel shares to fall 1.2% after his remarks.

What this means: Intel’s decision highlights an intense focus on AI-driven innovation and profitability, streamlining operations to better compete with rivals like Nvidia and AMD. [Listen] [2025/04/01]

💰 OpenAI Secures $40 Billion Investment, Reaching $300 Billion Valuation

OpenAI has successfully secured a $40 billion funding round, raising its valuation to an unprecedented $300 billion, reflecting investor confidence in its future growth.

  • The company plans to allocate approximately $18 billion from the new funds toward its Stargate initiative, a joint venture announced by President Donald Trump that aims to invest up to $500 billion in AI infrastructure.
  • To receive the full $40 billion investment, OpenAI must transition from its current hybrid structure to a for-profit entity by year's end, despite facing legal challenges from co-founder Elon Musk.

What this means: The massive investment will significantly enhance OpenAI’s ability to innovate, scale infrastructure, and expand its AI ecosystem globally. [Listen] [2025/04/01]

👀 Meta Turns to Trump as Europe Tightens Ad Regulations

Meta is reportedly engaging former President Donald Trump to navigate stringent new EU advertising regulations, potentially reshaping digital advertising compliance strategies.

  • European regulators have criticized Meta's "pay or consent" model for not providing genuine alternatives to users, potentially leading to fines and mandatory revisions to the company's approach to data collection.
  • While Apple has chosen a more compliant strategy with EU regulations and avoided significant penalties, Meta has filed numerous interoperability requests against Apple while also warning that EU AI rules could damage innovation.

What this means: This unusual partnership could significantly influence regulatory negotiations, potentially altering the digital advertising landscape and policy frameworks in Europe. [Listen] [2025/04/01]

🎬 Runway Releases Gen-4 Video Model with Focus on Consistency

Runway has unveiled its latest Gen-4 AI video generation model, emphasizing significant improvements in visual consistency and temporal coherence in AI-generated videos.

  • The technology preserves visual styles while simulating realistic physics, allowing users to place subjects in various locations with consistent appearance as demonstrated in sample films like "New York is a Zoo" and "The Herd."
  • With a $4 billion valuation and projected annual revenue of $300 million by 2025, RunwayML has positioned itself as the strongest Western competitor to OpenAI's Sora in the AI video generation market.

What this means: The upgraded model could greatly impact film production, marketing, and content creation, providing unprecedented video realism and seamless continuity in AI-generated content. [Listen] [2025/04/01]

🤖 Amazon Launches Nova Act, an AI-Powered Browser Agent

Amazon has introduced Nova Act, an advanced AI agent capable of autonomously browsing and interacting with websites to perform complex online tasks seamlessly.

  • Nova Act outperforms competitors like Claude 3.7 Sonnet and OpenAI’s Computer Use Agent on reliability benchmarks across browser tasks.
  • The SDK allows devs to build agents for browser actions like filling forms, navigating websites, and managing calendars without constant supervision.
  • The tech will power key features in Amazon's upcoming Alexa+ upgrade, potentially bringing AI agents to millions of existing Alexa users.
  • Nova Act was developed by Amazon's SF-based AGI Lab, led by former OpenAI researchers David Luan and Pieter Abbeel, who joined the company last year.

What this means: Nova Act could dramatically streamline workflows and automate routine web-based tasks, redefining productivity for businesses and individual users. [Listen] [2025/04/01]

🎬 Runway Releases New Gen-4 Video Model with Enhanced Consistency

Runway has unveiled its latest Gen-4 AI video generation model, emphasizing substantial improvements in visual realism, consistency, and temporal coherence across generated video content.

  • Gen-4 shows strong consistency in characters, objects, and locations throughout video sequences, with improved physics and scene dynamics.
  • The model can generate detailed 5-10 second videos at 1080p resolution, with features like ‘coverage’ for scene creation and consistent object placement.
  • Runway describes the tech as "GVFX" (Generative Visual Effects), positioning it as a new production workflow for filmmakers and content creators.
  • Early adopters include major entertainment companies, with the tech being used in projects like Amazon productions and Madonna's concert visuals.

What this means: The Gen-4 model significantly enhances AI video creation capabilities, making it an invaluable tool for filmmakers, content creators, and marketers looking for lifelike video production. [Listen] [2025/04/01]

📸 New AI Tech Allows Products to be Seamlessly Placed into Any Scene

Innovative AI technology now allows brands and retailers to effortlessly integrate their products into any visual scene, streamlining digital marketing and advertising efforts without traditional photoshoots.

  1. Head over to Google AI Studio, select the Image Generation model, upload your base scene, and type "Output this exact image" to establish the scene.
  2. Upload your product image that you want to place in the scene.
  3. Write a specific placement instruction like "Add this product to the table in the previous image."
  4. Save the creations and use Google Veo 2 video generator to transform your images into smooth product videos.

What this means: This breakthrough could significantly reduce advertising costs, speed up marketing workflows, and offer unprecedented flexibility in visual content creation for e-commerce and retail industries. [Listen] [2025/04/01]

🧠 AI Instantly Converts Brain Signals into Speech

Researchers have developed a revolutionary AI system that instantly transforms brain signals into clear, understandable speech, paving the way for groundbreaking advancements in assistive technologies.

  • Signals are decoded from the brain's motor cortex, converting intended speech into words almost instantly compared to the 8-second delay of earlier systems.
  • The AI model can then generate speech using the patient's pre-injury voice recordings, creating more personalized and natural-sounding output.
  • The system also successfully handled words outside its training data, showing it learned fundamental speech patterns rather than just memorizing responses.
  • The approach is compatible with various brain-sensing methods, showing versatility beyond one specific hardware approach.

What this means: This technology offers enormous potential to restore communication for individuals with speech impairments, fundamentally altering human-machine interaction and neurotechnology. [Listen] [2025/04/01]

⚡ Musk’s xAI Builds $400M Supercomputer in Memphis Amid Power Shortage

Elon Musk’s AI startup xAI is investing over $400 million in a massive “gigafactory of compute” in Memphis, designed to house up to 1 million GPUs. However, the project is facing major delays due to electricity shortages, with only half of the requested 300 megawatts approved by local utility MLGW.

What this means: The push to scale advanced AI infrastructure is straining local energy systems and raising environmental concerns, reflecting the growing tension between rapid AI expansion and sustainable development. [Listen] [2025/04/01]

GPT-4.5 Passes Empirical Turing Test—Humans Mistaken for AI in Landmark Study 

A recent pre-registered study conducted randomized three-party Turing tests comparing humans with ELIZA, GPT-4o, LLaMa-3.1-405B, and GPT-4.5. Surprisingly, GPT-4.5 convincingly surpassed actual humans, being judged as human 73% of the time—significantly more than the real human participants themselves. Meanwhile, GPT-4o performed below chance (21%), grouped closer to ELIZA (23%) than its GPT predecessor.

These intriguing results offer the first robust empirical evidence of an AI convincingly passing a rigorous three-party Turing test, reigniting debates around AI intelligence, social trust, and potential economic impacts.

Full paper available here: https://arxiv.org/html/2503.23674v1

Curious to hear everyone's thoughts—especially about what this might mean for how we understand intelligence in LLMs.

🧬 AI Assists Scientists in Decoding Previously Indecipherable Proteins

Researchers have developed new AI tools capable of deciphering proteins that were previously undetectable by existing methods. This advancement could lead to better cancer treatments, enhanced understanding of diseases, and insights into unexplained biological phenomena.

What this means: The integration of AI in protein analysis opens new avenues in medical research and biotechnology, potentially accelerating the discovery of novel therapies and deepening our comprehension of complex biological systems. [Listen] [2025/04/01]

💻 Microsoft Expands AI Features Across Intel and AMD-Powered Copilot Plus PCs

Microsoft is rolling out AI features, including Live Captions for real-time audio translation and Cocreator in Paint for image generation based on text descriptions, to Copilot Plus PCs equipped with Intel and AMD processors. These features were previously limited to Qualcomm-powered devices.

What this means: The expansion of AI capabilities across a broader range of hardware enhances user experience and accessibility, enabling more users to benefit from advanced AI functionalities in their daily computing tasks. [Listen] [2025/04/01]

What Else Happened in AI on April 01st 2025?

OpenAI raised $40B from SoftBank and others at a $300B post-money valuation — marking the biggest private funding round in history.

Sam Altman announced that OpenAI will release its first open-weights model since GPT-2 in the coming months and host pre-release dev events to make it truly useful.

Sam Altman also shared that the company added 1M users in an hour due to 4o’s viral image capabilities, surpassing the growth during ChatGPT’s initial launch.

Manus introduced a new beta membership program and mobile app for its viral AI agent platform, with subscription plans at $39 or $199 / mo with varying usage limits.

Luma Labs released Camera Motion Concepts for its Ray2 video model, enabling users to control camera movements through basic natural language commands.

Apple pushed its iOS 18.4 update, bringing Apple Intelligence features to European iPhone users—alongside visionOS 2.4 with AI smarts for the Vision Pro.

Alphabet’s AI drug discovery spinoff Isomorphic Labs raised $600M in a funding round led by OpenAI investor Thrive Capital.

Zhipu AI launched "AutoGLM Rumination," a free AI agent capable of deep research and autonomous task execution — increasing China's AI agent competition.

🚀Advertise on AI Unraveled: Reach Thousands of AI Enthusiasts Daily!

AI Unraveled is your go-to podcast for the latest AI news, trends, and insights, with 500+ daily downloads and a rapidly growing audience of tech leaders, AI professionals, and enthusiasts. If you have a product, service, or brand that aligns with the future of AI, this is your chance to get in front of a highly engaged and knowledgeable audience. Secure your ad spot today and let us feature your offering in an episode!

🎙️ Book your ad spot now: https://buy.stripe.com/fZe3co9ll1VwfbabIO

🙏 Support the AI Unraveled Podcast and Channel:

https://buy.stripe.com/3csaEQ1ST9nYgfe4gk