GeoFix: Geometry Fixing for 4D Consistent Generation

1KAIST    2ETH Zurich    3Google    4EPFL
†Corresponding authors

TL;DR

GeoFix aligns a video diffusion model with a single geometric reward, FOCAL (FOV-CALibrated geometric reward), computed directly on video latents and usable for both post-training and inference-time steering.

Key Contributions

  • A single geometric reward. FOCAL measures cross-view depth consistency in static scene regions, targeting scene structure without requiring colors or textures to match or penalizing valid object motion.
  • Better geometry, better video quality. Since geometric inconsistencies are themselves visual defects, fixing them with FOCAL alone improves WorldScore 3D consistency by 39% and VBench Quality Score by 4%, without any auxiliary appearance reward.
  • FOV calibration against overlap bias. Depth agreement correlates with camera overlap even in real videos and 4D datasets. FOCAL calibrates against this trend, so the model cannot score higher simply by moving the camera less.
  • Efficiency and versatility. FOCAL is computed from latents through a stitched 4D reconstructor, without VAE decoding. GeoFix applies to any reward alignment method without modification, which we demonstrate across training and inference-time alignment.

Method

Overview of GeoFix. (a) A 4D reconstructor (NeoVerse[1]) stitched to the video latent space[2] predicts 4D Gaussians, depth, and cameras directly from latents, without VAE decoding. (b) FOCAL measures cross-view depth reprojection error in static regions, so valid object motion is not penalized, and (c) calibrates it against the depth agreement expected for the camera overlap.

Observation: Overlap Bias

  • Depth agreement depends on the camera, not only on geometry. Even real videos and 4D datasets score higher when the first and last cameras' fields of view overlap more.
  • Existing rewards share the bias. Direct rewards (depth or RGB reprojection) and indirect ones (PickScore, VBench) all rise as the camera moves less.
  • A biased reward can be gamed. These scores cannot disentangle geometric quality from camera overlap, so maximizing them favors larger camera overlap without improvement in scene geometry.

FOCAL: FOV-Calibrated Reward

  • Calibrate on real videos and 4D datasets. We fit how the raw depth score \(R_{\mathrm{raw}}\) grows with FOV overlap \(\Omega\) on real videos and 4D datasets, once, and keep the fit fixed during alignment.
  • Rewards geometry, not camera choice. The calibrated score stays flat across camera overlap on held-out real and generated videos, so a higher FOCAL reflects better geometry rather than a smaller camera motion.

Text-to-4D Generation

Fine-tuning Wan2.2-TI2V-5B with FOCAL as the only reward yields cleaner 4D scenes than baselines, with substantially less background drift and flickering.

BaseSFTVIST3A[2]World-R1[3]GeoFix (Ours)

"The handheld camera follows the subject closely, simulating human perspective: 'Yellow school bus stopping in front of a red building'."

"The camera performs a smooth 360° panoramic rotation around the scene: 'Girl in pink sweater holding a golden trophy'. The motion fully encircles the environment."

"Pan the camera horizontally to uncover the subject and background in a fluid movement: 'A red fox sneaks past a green statue in the garden'."

"An artist in a paint-splattered apron demonstrates techniques to a boy in a t-shirt and jeans, both engaged in a creative workshop."

4D Scene Generation

The stitched model reconstructs video latents directly into an explicit 4D Gaussian scene. To keep the viewer light, we only leave a subset (about 30K Gaussians). 3D representations from multiple timesteps are overlaid. Select a scene below.

Loading 4DGS…
1 / 41

Comparison with Image-to-4D Methods

Compared with MoGe4D[4] and 4DNeX[5], GeoFix keeps the scene structure stable under novel camera trajectories while maintaining meaningful dynamics.

Qualitative comparison on image-to-4D generation with overlaid 3D representations over time

References

[1] Yang et al., NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos.

[2] Go et al., Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator (VIST3A).

[3] Wang et al., World-R1: Reinforcing 3D Constraints for Text-to-Video Generation.

[4] Zhang et al., Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation (MoGe4D).

[5] Chen et al., 4DNeX: Feed-Forward 4D Generative Modeling Made Easy.

BibTeX