Application (pre-grant publication)
TRAINING FOR MULTIMODAL CONDITIONAL 3D SHAPE GEOMETRY GENERATION
- Number
- 20250356582
- Published
- 2025-11-20
- Filed
- 2025-05-19
- Assignee
- DISNEY ENTERPRISES, INC.
- Inventors
- OTTO; Christopher Andreas et al.
- CPC
- G06N20/00; G06N3/045; G06N3/084; G06T15/205; G06T17/00; G06T17/10; G06T17/20; G06T19/20; G06T5/60
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Training technique for multimodal 3D-shape geometry generation (companion patent).
Abstract
One embodiment of the present invention sets forth a technique for training a machine learning model on a geometry generation task. The technique includes generating, via execution of a diffusion model, a first set of training output corresponding to a first set of three-dimensional (3D) geometries based on a first set of conditioning inputs associated with a first conditioning mode, and training the diffusion model based on a first set of loss values associated with the first set of training output. The technique further includes generating, via execution of the diffusion model and a first adapter model, a second set of training output corresponding to a second set of 3D geometries based on a second set of conditioning inputs associated with a second conditioning mode, and training the first adapter model based on a second set of loss values associated with the second set of training output.
Background
BACKGROUND Field of the Various Embodiments
Embodiments of the present disclosure relate generally machine learning and computer vision and, more specifically, to training for multimodal conditional three-dimensional (3D) shape geometry generation. Description of the Related Art
Realistic digital representations of faces, hands, bodies, and other recognizable objects are required for various computer graphics and computer vision applications. For example, digital representations of real-world deformable objects may be used in virtual scenes of film or television productions, video games, virtual worlds, and/or other environments and/or settings.
Traditionally, three-dimensional (3D) geometries of faces and/or other types of deformable objects have been generated via a time-consuming, iterative, and resource-intensive process involving digital sculpting with 3D modeling tools. For example, a user may spend days to weeks interacting with a 3D modeling tool to manually push, pull, smooth, grab, pinch, and/or otherwise manipulate a 3D geometry of a face. As the user interacts with the 3D geometry, the 3D modeling tool expends significant resources in updating a mesh and/or another 3D representation of the face based on sculpting input from the user, rendering the face to reflect the sculpting input, and/or outputting the rendered face to the user.
To simplify the task of modeling the 3D geometry of a face (or another type of deformable object), a param