Outer Rim Archives
Archives · 2025 · 20250356506

Application (pre-grant publication)

SEMANTIC VIDEO MOTION TRANSFER USING MOTION-TEXTUAL INVERSION

Number
20250356506
Published
2025-11-20
Filed
2025-05-16
Assignee
DISNEY ENTERPRISES, INC.
Inventors
KANSY; Manuel Jakob et al.
CPC
G06T7/246; G06T5/60; G06T5/70; G06T13/40; G06T13/80
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Semantic video motion-transfer VFX technique.

Abstract

One embodiment of the present invention sets forth a technique for performing motion transfer. The technique includes determining an embedding corresponding to a motion depicted in a first video. The technique also includes generating, via execution of a machine learning model based on the embedding and an appearance image, an output video that includes the motion depicted in the first video and an appearance depicted in the appearance image.

Background

BACKGROUND Field of the Various Embodiments

The present invention relates generally to computer vision and machine learning and, more specifically, to semantic video motion transfer using motion-textual inversion. DESCRIPTION OF THE RELATED ART

Recent developments in machine learning and computer vision have led to significant improvements in the quality and functionality of video generation and editing techniques. For example, a diffusion model, which operates by iteratively converting random noise into new data such as images, can be trained to synthesize spatially and temporally coherent sequences of video frames. The diffusion model may operate as an image-to-video model that uses an image that acts as a starting or conditioning frame for the generation of the video and/or as a text-to-video model that uses a natural language description as input to produce a corresponding video. The diffusion model can also, or instead, be used to change the content, background, motion, and/or other attributes of an input video based on an input text prompt.

Existing techniques for generating and editing videos are typically unable to control both the appearance and motion in a video in a predictable and/or fine-grained manner. More specifically, the motion in a video generated by a conventional image-to-video diffusion model may be modified by altering the random seed used to generate random noise that is converted into the video and/or adjusting micro-conditioning

Claims

1. A computer-implemented method for performing motion transfer, the method comprising: determining an embedding corresponding to a motion depicted in a first video; and generating, via execution of a machine learning model based on the embedding and an appearance image, an output video that includes the motion depicted in the first video and an appearance depicted in the appearance image. || 11. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: determining an embedding corresponding to a motion depicted in a first video; and generating, via execution of a machine learning model based on the embedding and an appearance image, an output video that includes the motion depicted in the first video and an appearance depicted in the appearance image. || 20. A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to perform the steps of: determining an embedding corresponding to a motion depicted in a first video; and generating, via execution of a machine learning model based on the embedding, an output video that includes the motion depicted in the first video and an appearance that is not depicted in the first video.