Generative Cinematographer:
Composing Camera and Object Motion in 3D

Jiahan Zhang1, Chaohao Yang1, Namitha Guruprasad1, Vivekjyoti Banerjee1, Trong-Tung Nguyen1, Alan Yuille1, Anand Bhattad1

1Johns Hopkins University

Example:
Input image
Input image for the selected scene
Authored 3D controls
Generated video

Camera and object motion are authored together in 3D from a single image. Orange: camera. Other colors: object paths.

Overview

TL;DR

Author the camera path and object motion together in a 3D scene lifted from a single image, and a video diffusion model turns them into a coherent shot.

Contributions

  1. Correspondence maps that encode each point's world position and identity as images a pretrained video VAE can read.
  2. Local 3D motion handles: independently moved rigid regions that approximate non-rigid motion, with no simulator or category prior.
  3. A world-space interface for authoring camera paths and object motion together from one image.

Method

Framework overview. Click to enlarge, or open the PDF.
  1. Lift the image into a colored point cloud using estimated depth and a dynamic mask.
  2. Select foreground regions as motion handles in a 3D tool such as Blender, then author their paths and the camera path.
  3. Render each frame into three maps: background XYZ, foreground XYZ, and foreground identity.
  4. Encode the maps with the frozen Wan VAE, and condition the frozen diffusion model through a side branch and LoRA adapters.

Only the side branch, its projections, and the LoRA adapters are trained. The side branch runs at half resolution and feeds the main branch through residual connections.

Correspondence maps

Each XYZ map stores a point's world position as RGB, normalized by the first-frame bounding box (center o, half-range σ) in every frame:

RGB = ½ [(p − o) / σ + 1]

A pixel's position shows where a point appears; its color shows where it is in the world. Moving points change color, so an identity map gives each handle a fixed color.

3D scene (fixed overview)

Static backgroundMotion handleCamera

Background XYZ

Static points keep the same colors in every view.

World reference

Foreground XYZ

The color changes when the handle moves.

RGB

Foreground identity

A fixed color that labels the handle.

Handle 1

Move the camera: points shift in the image, but their colors stay the same.

Background points keep their colors across viewpoints while the bus moves. PDF.

Comparisons

Pick a scene and a baseline. The left panel shows the control input.

Camera overview
Ours

Orange: camera. Other colors: object paths.

All examples are hand-authored. Baselines use different control inputs.

Results

Limitations

Single-view geometry is incomplete. Depth errors and hidden surfaces can cause artifacts under large viewpoint changes.

Handles are locally rigid. Tearing, fracture, and fluids are out of scope, and fine deformation needs many handles.

Trajectories are not physically constrained, and large edits can exceed the first-frame XYZ range.

BibTeX

@misc{zhang2026generativecinematographer,
  title  = {Generative Cinematographer: Composing Camera and Object Motion in 3D},
  author = {Zhang, Jiahan and Yang, Chaohao and Guruprasad, Namitha and Banerjee, Vivekjyoti
            and Nguyen, Trong-Tung and Yuille, Alan and Bhattad, Anand},
  year   = {2026},
  eprint = {2610.02180},
  archivePrefix = {arXiv}
}