Outer Rim Archives
Archives · 2026 · 20260212267

Application (pre-grant publication)

METHOD AND SYSTEM FOR GENERATING TRAINING DATA FOR EFFICIENT AND ACCURATE TRAINING OF NOVEL VIEW SYNTHESIS MODELS

Number
20260212267
Published
2026-07-23
Filed
2025-07-15
Assignee
Disney Enterprises, Inc.
Inventors
Zhang; Xiang, Jiao; Longxiang, Zhang; Yang, Schroers; Christopher Richard, Mehl; Lukas Francesco
CPC
G06N3/044; G06T15/205; G06N3/045; G06N3/0455; G06N3/0464; G06N3/047; G06N3/08; G06N3/084; G06N3/088; G06N3/0895; G06N20/00; G06T3/18; G06T7/12; G06T7/13; G06T7/20; G06T7/50; G06T9/00; G06T11/00; G06T11/60; G06T15/20; G06V10/774; G06V10/806; G06T2207/20084
Verdict
Medium Notable software
First reported
2026-W30 (2026-07-24)
Source
Google Patents · FreePatentsOnline

The keeper's note

Methods and systems for performing novel view synthesis in addition to various image-to-image diffusion processes using machine learning models are disclosed.

Abstract

Methods and systems for performing novel view synthesis in addition to various image-to-image diffusion processes using machine learning models are disclosed. These include methods for performing inpainting-based novel view synthesis using depth maps and conditional models, video novel view synthesis using texture map inpainting, spatially consistent inpainting-based novel view synthesis using video diffusion models, and high-fidelity image-to-image diffusion using a texture conditional model. Further, methods for generating training data that can be used to train machine learning models to perform novel view synthesis are disclosed. These include symmetry exploiting methods of training data generation and methods using training pair alignment and splatting error simulations.

Background

BACKGROUND

The field of novel view synthesis (NVS) generally involves using input data to produce a new view of a scene, e.g., comprising a digital image or video of that scene. As an example, NVS can be used to generate a digital image showing what a scene would look like if an imaginary camera (corresponding to that scene) was moved.

Novel view synthesis can be used in various tasks, including the generation of stereo images and videos. Recent developments in virtual reality (VR) technology, including an increase in commercially available VR headsets, have led to increased interest in stereo images and stereo videos. In a stereo image (or video), two images (or videos) corresponding to slightly different camera angles are displayed to two eyes independently (e.g., from two screens located on the inside of a VR headset). The disparity between the two images or videos corresponds to the apparent displacement of objects in those images or videos when viewed from one eye or the other, which is related to the apparent distance between the objects and the observer. This is similar to the disparity of visual information that people receive when viewing objects with two eyes. As such, viewing images or videos in this manner leads to a “stereoscopic viewing effect,” allowing the brain interpret images or videos as three dimensional (3D). This can be desirable in various applications, including VR games or other media, as it can make the viewer feel like they are “immer

Claims

1. A method for training a machine learning model to generate novel view images corresponding to warped images using a single-view dataset, the method performed by a computer system and comprising: retrieving the single-view dataset, wherein the single-view dataset comprises a plurality of images; for each image: performing depth estimation to generate a depth map corresponding to the image, generating an occlusion mask based on a warping process and the depth map, generating a masked image by masking the image with the occlusion mask, and generating a training set of images comprising the image and the masked image, thereby generating a training set of images for each image in the single-view dataset, thereby generating a plurality of training sets of images; and training the machine learning model using the plurality of training sets of images. || 12. A method for training machine learning model to generate novel view images corresponding to warped images, the method performed by a computer system and comprising: retrieving a plurality of sets of images; for each set of images: determining an input view image and a target view image from the set of images, determining an optical flow map by performing an optical flow estimation process between the input view image and the target view image, determining a warping mask based on the optical flow map, masking the target view image using the warping mask, thereby generating a masked target view image, generating a training set of images comprising the target view image and the masked target view image, thereby generating a training set of images for each set of images, thereby generating a plurality of training sets of images; and training the machine learning model using the plurality of training sets of images. || 20. A computer system comprising: one or more processors; and a non-transitory computer readable medium coupled to the one or more processors, the non-transitory computer readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform a method for training a machine learning model to generate novel view images corresponding to warped images using a single-view dataset, the method comprising: retrieving the single-view dataset, wherein the single-view dataset comprises a plurality of images; for each image: performing depth estimation to generate a depth map corresponding to the image, generating an occlusion mask based on a warping process and the depth map; generating a masked image by masking the image with the occlusion mask, and generating a training set of images comprising the image and the masked image, thereby generating a training set of images for each image in the single-view dataset, thereby generating a plurality of training sets of images; and training the machine learning model using the plurality of training sets of images.