Outer Rim Archives
Archives · 2026 · 20260212590

Application (pre-grant publication)

VISUALLY CONSISTENT MULTI-VIEW NOVEL VIEW SYNTHESIS WITH VIDEO DIFFUSION

Number
20260212590
Published
2026-07-23
Filed
2025-07-15
Assignee
Disney Enterprises, Inc.
Inventors
Schroers; Christopher Richard, Zhang; Yang, Zhang; Xiang, Mehl; Lukas Francesco
CPC
G06N3/044; G06T15/205; G06N3/045; G06N3/0455; G06N3/0464; G06N3/047; G06N3/08; G06N3/084; G06N3/088; G06N3/0895; G06N20/00; G06T3/18; G06T7/12; G06T7/13; G06T7/20; G06T7/50; G06T9/00; G06T11/00; G06T11/60; G06T15/20; G06V10/774; G06V10/806; G06T2207/20084
Verdict
Medium Notable software
First reported
2026-W30 (2026-07-24)
Source
Google Patents · FreePatentsOnline

The keeper's note

Methods and systems for performing novel view synthesis in addition to various image-to-image diffusion processes using machine learning models are disclosed.

Abstract

Methods and systems for performing novel view synthesis in addition to various image-to-image diffusion processes using machine learning models are disclosed. These include methods for performing inpainting-based novel view synthesis using depth maps and conditional models, video novel view synthesis using texture map inpainting, spatially consistent inpainting-based novel view synthesis using video diffusion models, and high-fidelity image-to-image diffusion using a texture conditional model. Further, methods for generating training data that can be used to train machine learning models to perform novel view synthesis are disclosed. These include symmetry exploiting methods of training data generation and methods using training pair alignment and splatting error simulations.

Background

BACKGROUND

The field of novel view synthesis (NVS) generally involves using input data to produce a new view of a scene, e.g., comprising a digital image or video of that scene. As an example, NVS can be used to generate a digital image showing what a scene would look like if an imaginary camera (corresponding to that scene) was moved.

Novel view synthesis can be used in various tasks, including the generation of stereo images and videos. Recent developments in virtual reality (VR) technology, including an increase in commercially available VR headsets, have led to increased interest in stereo images and stereo videos. In a stereo image (or video), two images (or videos) corresponding to slightly different camera angles are displayed to two eyes independently (e.g., from two screens located on the inside of a VR headset). The disparity between the two images or videos corresponds to the apparent displacement of objects in those images or videos when viewed from one eye or the other, which is related to the apparent distance between the objects and the observer. This is similar to the disparity of visual information that people receive when viewing objects with two eyes. As such, viewing images or videos in this manner leads to a “stereoscopic viewing effect,” allowing the brain interpret images or videos as three dimensional (3D). This can be desirable in various applications, including VR games or other media, as it can make the viewer feel like they are “immer

Claims

1. A method for generating a plurality of novel view images corresponding to an input view image using a video diffusion model, the method performed by a computer system and comprising: generating, based on the input view image, a depth map; warping the input view image using the depth map, thereby generating a plurality of warped images; encoding the plurality of warped images, thereby generating a plurality of warped image embeddings; generating one or more noisy embeddings; applying the plurality of warped image embeddings and the one or more noisy embeddings to the video diffusion model, thereby generating a plurality of output embeddings; and decoding the plurality of output embeddings, thereby generating the plurality of novel view images. || 11. A method for generating one or more output images based on one or more input images using a machine learning model comprising an encoder, a diffusion model, a decoder, and a texture conditional model, and wherein the encoder comprises one or more encoder layers, the decoder comprises one or more decoder layers, the method performed by a computer system and comprising: encoding the one or more input images, thereby generating one or more input image embeddings and one or more sets of encoder features, wherein each set of encoder features corresponds to an input image of the one or more input images, wherein each set of encoder features comprises one or more encoder features corresponding to the one or more encoder layers; generating one or more noisy embeddings; combining the one or more input image embeddings and the one or more noisy embeddings, thereby generating one or more combined embeddings; applying the one or more combined embeddings to the diffusion model, thereby generating one or more output embeddings; applying the one or more output embeddings to the decoder, thereby generating one or more sets of decoder features, wherein each set of decoder features corresponds to an input image of the one or more input images, wherein each set of decoder features comprises one or more decoder features corresponding to the one or more decoder layers; combining each set of encoder features and a corresponding set of decoder features using the texture conditional model, thereby generating one or more sets of fused features; and applying the one or more sets of fused features to the one or more decoder layers, thereby generating one or more output images. || 20. A computer system comprising: one or more processors; and a non-transitory computer readable medium coupled to the one or more processors, the non-transitory computer readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform a method for generating a plurality of novel view images corresponding to an input view image using a video diffusion model, the method performed by a computer system and comprising: generating, based on the input view image, a depth map; warping the input view image using the depth map, thereby generating a plurality of warped images; encoding the plurality of warped images, thereby generating a plurality of warped image embeddings; generating one or more noisy embeddings; applying the plurality of warped image embeddings and the one or more noisy embeddings to the video diffusion model, thereby generating a plurality of output embeddings; and decoding the plurality of output embeddings, thereby generating the plurality of novel view images.