Application (pre-grant publication)
SEMANTIC VIDEO MOTION TRANSFER USING MOTION-TEXTUAL INVERSION
- Number
- 20250356506
- Published
- 2025-11-20
- Filed
- 2025-05-16
- Assignee
- DISNEY ENTERPRISES, INC.
- Inventors
- KANSY; Manuel Jakob et al.
- CPC
- G06T7/246; G06T5/60; G06T5/70; G06T13/40; G06T13/80
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Semantic video motion-transfer VFX technique.
Abstract
One embodiment of the present invention sets forth a technique for performing motion transfer. The technique includes determining an embedding corresponding to a motion depicted in a first video. The technique also includes generating, via execution of a machine learning model based on the embedding and an appearance image, an output video that includes the motion depicted in the first video and an appearance depicted in the appearance image.
Background
BACKGROUND Field of the Various Embodiments
The present invention relates generally to computer vision and machine learning and, more specifically, to semantic video motion transfer using motion-textual inversion. DESCRIPTION OF THE RELATED ART
Recent developments in machine learning and computer vision have led to significant improvements in the quality and functionality of video generation and editing techniques. For example, a diffusion model, which operates by iteratively converting random noise into new data such as images, can be trained to synthesize spatially and temporally coherent sequences of video frames. The diffusion model may operate as an image-to-video model that uses an image that acts as a starting or conditioning frame for the generation of the video and/or as a text-to-video model that uses a natural language description as input to produce a corresponding video. The diffusion model can also, or instead, be used to change the content, background, motion, and/or other attributes of an input video based on an input text prompt.
Existing techniques for generating and editing videos are typically unable to control both the appearance and motion in a video in a predictable and/or fine-grained manner. More specifically, the motion in a video generated by a conventional image-to-video diffusion model may be modified by altering the random seed used to generate random noise that is converted into the video and/or adjusting micro-conditioning