Outer Rim Archives
Archives · 2025 · 20250355935

Application (pre-grant publication)

GAUSSIAN SPLATTING WITH NEURAL SPLINE DEFORMATION

Number
20250355935
Published
2025-11-20
Filed
2025-05-16
Assignee
DISNEY ENTERPRISES, INC.
Inventors
ZHANG; Yang et al.
CPC
G06F16/735; G06F16/738; G06F16/7867; G06V10/75; G06V10/82
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Gaussian splatting + neural spline deformation rendering technique.

Abstract

One embodiment of the present invention sets forth a technique for determining a time-varying deformation associated with a scene. The technique includes matching a query time to a time interval associated with the scene and generating, via execution of a machine learning model, (i) a first set of attributes associated with a set of canonical coordinates in the scene at a starting time of the time interval and (ii) a second set of attributes associated with the set of canonical coordinates at an ending time of the time interval. The technique also includes computing a third set of attributes associated with the set of canonical coordinates at the query time based on a spline interpolation associated with the first and second sets of attributes. The technique further includes generating a representation of the scene at the query time based on the third set of attributes.

Background

BACKGROUND Field of the Various Embodiments

Embodiments of the present disclosure relate generally to machine learning and computer vision and, more specifically, to Gaussian splatting with neural spline deformation. DESCRIPTION OF THE RELATED ART

Films, video games, virtual reality (VR) systems, augmented reality (AR) systems, mixed reality (MR) systems, motion capture, and/or other types of applications frequently involve generating and/or making changes to depictions of 3D scenes over time. Traditionally, a visual representation of a given scene is generated and/or edited via a time-consuming, iterative, and/or laborious process. For example, a conventional visual effects workflow may involve a visual effects artist adding special effects and/or posing or animating a virtual character on a frame-by-frame basis.

More recently, advancements in machine learning and deep learning have led to the development of neural deformation models, which include deep neural networks that learn implicit representations of non-rigid and/or time-varying scenes. These neural deformation models commonly include coordinate neural networks that map coordinates in a canonical space to corresponding deformed coordinates at various temporal offsets. The deformed coordinates can then be used to render and/or reconstruct the corresponding scenes at the temporal offsets.

However, conventional neural deformation models are associated with a tradeoff between performance and a

Claims

1. A computer-implemented method for determining a time-varying deformation associated with a scene, the method comprising: matching a query time to a time interval associated with the scene; generating, via execution of a machine learning model, (i) a first set of attributes associated with a set of canonical coordinates in the scene at a starting time of the time interval and (ii) a second set of attributes associated with the set of canonical coordinates at an ending time of the time interval; computing a third set of attributes associated with the set of canonical coordinates at the query time based on a spline interpolation associated with the first set of attributes and the second set of attributes; and generating a representation of the scene at the query time based on the third set of attributes. || 11. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: matching a query time to a time interval associated with a scene; generating, via execution of a machine learning model, (i) a first set of attributes associated with a set of canonical coordinates of a three-dimensional (3D) Gaussian in the scene at a starting time of the time interval and (ii) a second set of attributes associated with the set of canonical coordinates at an ending time of the time interval; computing a third set of attributes associated with the set of canonical coordinates at the query time based on a spline interpolation associated with the first set of attributes and the second set of attributes; and generating a representation of the scene at the query time based on the third set of attributes. || 20. A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to perform the steps of: matching a query time to a time interval associated with a scene; generating, via execution of a machine learning model based on the query time and a set of canonical coordinates of a three-dimensional (3D) Gaussian in the scene, (i) a first set of deformed coordinates at a starting time of the time interval and (ii) a second set of deformed coordinates at an ending time of the time interval; computing a third set of deformed coordinates at the query time based on a spline interpolation associated with the first set of deformed coordinates and the second set of deformed coordinates; and generating a representation of the scene at the query time based on the third set of deformed coordinates.