Outer Rim Archives
Archives · 2024 · 20240312035

Application (pre-grant publication)

REAL-TIME MATTING SYSTEM

Number
20240312035
Published
2024-09-19
Filed
2023-03-16
Assignee
DISNEY ENTERPRISES, INC.
Inventors
Guay; Martin et al.
CPC
G06T7/50; G06V10/70; G06T5/50; G06T7/194
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Real-time image-matting VFX technique.

Abstract

The disclosed matting technique comprises receiving a video feed comprising a plurality of temporally ordered video frames as the video frames are captured by a video capture device, generating, using one or more machine learning models, an image mask corresponding to each video frame included in the video feed and a depth estimate corresponding to each video frame included in the video feed, and, for each video frame in the video feed, transmitting, in real-time, the video frame, the corresponding image mask, and the corresponding depth estimate to a compositing system. The compositing system composites a computer-generated element with the video frame based on the corresponding image mask and the corresponding depth estimate.

Background

BACKGROUND Field of the Various Embodiments

Embodiments of the present disclosure relate generally to computer science and video processing and, more specifically, to a real-time matting system. Description of the Related Art

Oftentimes, live-action video footage of an actor needs to be composed with one or more computer-generated elements, such as a virtual character, virtual object, background, and/or the like. For example, the background of the live-action video footage could be replaced with a computer-generated background. As another example, a computer-generated character could be added to a scene such that the actor appears to interact with the computer-generated character.

One approach for composing live-action video footage with computer-generated elements is to use a green screen or motion capture volumes. However, green screens and motion capture both require expensive equipment and complex setups. Additionally, processing live-action video footage that utilizes green screen and/or motion capture requires significant processing power and time. Accordingly, using these approaches, a composite video cannot be quickly reviewed in order to make changes to the live-action footage while filming.

Another approach is to utilize software applications that can add computer-generated elements to video or images captured on a personal device, such as a mobile phone or computer webcam. However, the software applications are usually only able to modif

Claims

1. A computer-implemented method for compositing video frames with one or more computer-generated elements, the method comprising: receiving a video feed comprising a plurality of temporally ordered video frames as the video frames are captured by a video capture device; generating, using one or more machine learning models, an image mask corresponding to each video frame included in the video feed and a depth estimate corresponding to each video frame included in the video feed; and for each video frame in the video feed, transmitting, in real-time, the video frame, the corresponding image mask, and the corresponding depth estimate to a compositing system, wherein the compositing system is configured to composite a computer-generated element with the video frame based on the corresponding image mask and the corresponding depth estimate. || 10. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors cause the one or more processors to perform the steps of: receiving a video feed comprising a plurality of temporally ordered video frames as the video frames are captured by a video capture device; generating, using one or more machine learning models, an image mask corresponding to each video frame included in the video feed and a depth estimate corresponding to each video frame included in the video feed; and for each video frame in the video feed, transmitting, in real-time, the video frame, the corresponding image mask, and the corresponding depth estimate to a compositing system, wherein the compositing system is configured to composite a computer-generated element with the video frame based on the corresponding image mask and the corresponding depth estimate. || 19. A system, comprising: a memory storing instructions; and one or more processors to: receive a video feed comprising a plurality of temporally ordered video frames as the video frames are captured by a video capture device; generate, using one or more machine learning models, an image mask corresponding to each video frame included in the video feed and a depth estimate corresponding to each video frame included in the video feed; and for each video frame in the video feed, transmitting, in real-time, the video frame, the corresponding image mask, and the corresponding depth estimate to a compositing system, wherein the compositing system is configured to composite a computer-generated element with the video frame based on the corresponding image mask and the corresponding depth estimate.