Application (pre-grant publication)
MULTIMODAL CONDITIONAL 3D SHAPE GEOMETRY GENERATION
- Number
- 20250356586
- Published
- 2025-11-20
- Filed
- 2025-05-19
- Assignee
- DISNEY ENTERPRISES, INC.
- Inventors
- OTTO; Christopher Andreas et al.
- CPC
- G06N20/00; G06N3/045; G06N3/084; G06T15/205; G06T17/00; G06T17/10; G06T17/20; G06T19/20; G06T5/60
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Multimodal generative 3D-shape geometry synthesis technique.
Abstract
One embodiment of the present invention sets forth a technique for generating a geometry for a shape. The technique includes inputting, into a machine learning model, (i) a noise sample and (ii) one or more conditioning inputs. The technique also includes generating, via execution of the machine learning model based on the noise sample and the one or more conditioning inputs, a two-dimensional (2D) position map associated with the shape. The technique further includes generating a three-dimensional (3D) geometry for the shape based on the 2D position map.
Background
BACKGROUND Field of the Various Embodiments
Embodiments of the present disclosure relate generally machine learning and computer vision and, more specifically, to multimodal conditional three-dimensional (3D) shape geometry generation. Description of the Related Art
Realistic digital representations of faces, hands, bodies, and other recognizable objects are required for various computer graphics and computer vision applications. For example, digital representations of real-world deformable objects may be used in virtual scenes of film or television productions, video games, virtual worlds, and/or other environments and/or settings.
Traditionally, three-dimensional (3D) geometries of faces and/or other types of deformable objects have been generated via a time-consuming, iterative, and resource-intensive process involving digital sculpting with 3D modeling tools. For example, a user may spend days to weeks interacting with a 3D modeling tool to manually push, pull, smooth, grab, pinch, and/or otherwise manipulate a 3D geometry of a face. As the user interacts with the 3D geometry, the 3D modeling tool expends significant resources in updating a mesh and/or another 3D representation of the face based on sculpting input from the user, rendering the face to reflect the sculpting input, and/or outputting the rendered face to the user.
To simplify the task of modeling the 3D geometry of a face (or another type of deformable object), a parametric shape m