Outer Rim Archives
Archives · 2025 · 20250329083

Application (pre-grant publication)

VISUAL DUBBING OF AN AUDIOVISUAL SEQUENCE

Number
20250329083
Published
2025-10-23
Filed
2024-04-22
Assignee
DISNEY ENTERPRISES, INC.
Inventors
NARUNIEC; Jacek Krzysztof et al.
CPC
G06V10/774; G06T9/00; G06T11/60; G06V40/165; G06V40/171
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Visual/face dubbing VFX technique.

Abstract

The present invention sets forth a technique for performing visual dubbing on an audiovisual sequence. The technique includes identifying, based on an actor frame included in the audiovisual sequence, one or more regions of an actor's face included in the actor frame, identifying, based on a dubber frame included in a visual recording of a dubber's performance, one or more regions of a dubber's face included in the dubber frame, generating a plurality of latent vectors based on at least one identified region of the actor's face and at least one identified region of the dubber's face, and generating, via the machine learning model, an output image based on the plurality of latent vectors.

Background

BACKGROUND Field of the Various Embodiments

Embodiments of the present disclosure relate generally to machine learning and video effects processing and, more specifically, to techniques for dubbing an audiovisual sequence. Description of the Related Art

During the production of a live action or animated audiovisual sequence, producers, creators, dubbing directors, or distributors may wish to dub or replace one or more lines of dialogue in the audiovisual sequence with an alternate audio recording. For example, producers, creators, dubbing directors, or distributors may wish to generate a localized version of the audiovisual sequence, where dialogue included in the audiovisual sequence is replaced with a translation of the dialogue into a different language. Producers, creators, dubbing directors, or distributors may also wish to replace dialogue with an alternate version, with or without translation, to correct errors in the spoken dialogue, to achieve a different artistic goal, or to comply with ratings guidelines or societal standards.

Existing techniques for dubbing audiovisual sequences may simply replace a section of the original audio included in the audiovisual sequence with an alternate audio recording. These techniques require only that the duration of the alternate audio recording approximately matches the duration of the original audio included in the audiovisual sequence. One drawback of these techniques is that the techniques do not perform a

Claims

1. A computer-implemented method for performing visual dubbing of an audiovisual sequence, the computer-implemented method comprising: identifying, based on an actor frame included in the audiovisual sequence, one or more regions included in the actor frame; identifying, based on a dubber frame included in a visual recording of a dubber performance, one or more regions included in the dubber frame; generating a plurality of latent vectors based on at least one identified region included in the actor frame and at least one identified region included in the dubber frame; and generating an output image based on the plurality of latent vectors. || 15. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: identifying, based on an actor frame included in an audiovisual sequence, one or more regions included in the actor frame; identifying, based on a dubber frame included in a visual recording of a dubber performance, one or more regions included in the dubber frame; generating a plurality of latent vectors based on at least one identified region included in the actor frame and at least one identified region included in the dubber frame; and generating an output image based on the plurality of latent vectors.