Archive republication: This article was originally published on 7 September 2020. It has been translated into English and lightly updated for clarity while preserving the original reporting and context.
World-Consistent Video-to-Video Synthesis
Nvidia researchers presented a method for generating photorealistic video from semantic inputs while maintaining consistency over time. The architecture used a multi-SPADE module, memory, optical-flow warping and guidance images to reduce temporal discontinuities.
The project page provided videos and diagrams, while the paper documented the method and evaluation. Results from research examples should be tested against representative production material, especially for temporal artefacts, identity preservation and controllability.
Read the project page and the research paper. Related archive items cover water simulation and cloud rendering.