Overview
TL;DR
Author the camera path and object motion together in a 3D scene lifted from a single image, and a video diffusion model turns them into a coherent shot.
Contributions
- Correspondence maps that encode each point's world position and identity as images a pretrained video VAE can read.
- Local 3D motion handles: independently moved rigid regions that approximate non-rigid motion, with no simulator or category prior.
- A world-space interface for authoring camera paths and object motion together from one image.
Method
- Lift the image into a colored point cloud using estimated depth and a dynamic mask.
- Select foreground regions as motion handles in a 3D tool such as Blender, then author their paths and the camera path.
- Render each frame into three maps: background XYZ, foreground XYZ, and foreground identity.
- Encode the maps with the frozen Wan VAE, and condition the frozen diffusion model through a side branch and LoRA adapters.
Only the side branch, its projections, and the LoRA adapters are trained. The side branch runs at half resolution and feeds the main branch through residual connections.
Correspondence maps
Each XYZ map stores a point's world position as RGB, normalized by the first-frame bounding box (center o, half-range σ) in every frame:
A pixel's position shows where a point appears; its color shows where it is in the world. Moving points change color, so an identity map gives each handle a fixed color.
3D scene (fixed overview)
Background XYZ
Static points keep the same colors in every view.
Foreground XYZ
The color changes when the handle moves.
Foreground identity
A fixed color that labels the handle.
Move the camera: points shift in the image, but their colors stay the same.
Comparisons
Pick a scene and a baseline. The left panel shows the control input.
Orange: camera. Other colors: object paths.
All examples are hand-authored. Baselines use different control inputs.
Results
Limitations
Single-view geometry is incomplete. Depth errors and hidden surfaces can cause artifacts under large viewpoint changes.
Handles are locally rigid. Tearing, fracture, and fluids are out of scope, and fine deformation needs many handles.
Trajectories are not physically constrained, and large edits can exceed the first-frame XYZ range.
BibTeX
@misc{zhang2026generativecinematographer,
title = {Generative Cinematographer: Composing Camera and Object Motion in 3D},
author = {Zhang, Jiahan and Yang, Chaohao and Guruprasad, Namitha and Banerjee, Vivekjyoti
and Nguyen, Trong-Tung and Yuille, Alan and Bhattad, Anand},
year = {2026},
eprint = {2610.02180},
archivePrefix = {arXiv}
}