Outer Rim Archives
Archives · 2025 · 20250356582

Application (pre-grant publication)

TRAINING FOR MULTIMODAL CONDITIONAL 3D SHAPE GEOMETRY GENERATION

Number
20250356582
Published
2025-11-20
Filed
2025-05-19
Assignee
DISNEY ENTERPRISES, INC.
Inventors
OTTO; Christopher Andreas et al.
CPC
G06N20/00; G06N3/045; G06N3/084; G06T15/205; G06T17/00; G06T17/10; G06T17/20; G06T19/20; G06T5/60
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Training technique for multimodal 3D-shape geometry generation (companion patent).

Abstract

One embodiment of the present invention sets forth a technique for training a machine learning model on a geometry generation task. The technique includes generating, via execution of a diffusion model, a first set of training output corresponding to a first set of three-dimensional (3D) geometries based on a first set of conditioning inputs associated with a first conditioning mode, and training the diffusion model based on a first set of loss values associated with the first set of training output. The technique further includes generating, via execution of the diffusion model and a first adapter model, a second set of training output corresponding to a second set of 3D geometries based on a second set of conditioning inputs associated with a second conditioning mode, and training the first adapter model based on a second set of loss values associated with the second set of training output.

Background

BACKGROUND Field of the Various Embodiments

Embodiments of the present disclosure relate generally machine learning and computer vision and, more specifically, to training for multimodal conditional three-dimensional (3D) shape geometry generation. Description of the Related Art

Realistic digital representations of faces, hands, bodies, and other recognizable objects are required for various computer graphics and computer vision applications. For example, digital representations of real-world deformable objects may be used in virtual scenes of film or television productions, video games, virtual worlds, and/or other environments and/or settings.

Traditionally, three-dimensional (3D) geometries of faces and/or other types of deformable objects have been generated via a time-consuming, iterative, and resource-intensive process involving digital sculpting with 3D modeling tools. For example, a user may spend days to weeks interacting with a 3D modeling tool to manually push, pull, smooth, grab, pinch, and/or otherwise manipulate a 3D geometry of a face. As the user interacts with the 3D geometry, the 3D modeling tool expends significant resources in updating a mesh and/or another 3D representation of the face based on sculpting input from the user, rendering the face to reflect the sculpting input, and/or outputting the rendered face to the user.

To simplify the task of modeling the 3D geometry of a face (or another type of deformable object), a param

Claims

1. A computer-implemented method for training a machine learning model on a geometry generation task, the method comprising: generating, via execution of a diffusion model, a first set of training output corresponding to a first set of three-dimensional (3D) geometries based on a first set of conditioning inputs associated with a first conditioning mode; training the diffusion model based on a first set of loss values associated with the first set of training output; generating, via execution of the diffusion model and a first adapter model, a second set of training output corresponding to a second set of 3D geometries based on a second set of conditioning inputs associated with a second conditioning mode; and training the first adapter model based on a second set of loss values associated with the second set of training output. || 11. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: generating, via execution of a diffusion model, a first set of training output corresponding to a first set of three-dimensional (3D) geometries based on a first set of conditioning inputs associated with a first conditioning mode; training the diffusion model based on a first set of loss values associated with the first set of training output; generating, via execution of the diffusion model and one or more adapter models, one or more additional sets of training output corresponding to one or more additional sets of 3D geometries based on one or more additional sets of conditioning inputs associated with one or more additional conditioning modes; and training the one or more adapter models based on one or more additional sets of loss values associated with the one or more additional sets of training output. || 20. A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to perform the steps of: generating, via execution of a diffusion model, a first set of training output corresponding to a first set of three-dimensional (3D) geometries for a deformable object based on a first set of conditioning inputs associated with a first conditioning mode; training the diffusion model based on a first set of loss values associated with the first set of training output; generating, via execution of the diffusion model and a first adapter model, a second set of training output corresponding to a second set of 3D geometries for the deformable object based on a second set of conditioning inputs associated with a second conditioning mode; and training the first adapter model based on a second set of loss values associated with the second set of training output.