Outer Rim Archives
Archives · 2023 · 20230077379

Application (pre-grant publication)

MACHINE LEARNING BASED VIDEO COMPRESSION

Number
20230077379
Published
2023-03-16
Filed
2022-10-24
Assignee
DISNEY ENTERPRISES, INC.
Inventors
Schroers; Christopher et al.
CPC
H04N19/503; H04N19/54; H04N19/587; H04N19/436; H04N19/537
Verdict
Set aside video compression codec, plumbing
Source
Google Patents · FreePatentsOnline

Abstract

Systems and methods are disclosed for compressing a target video. A computer-implemented method may use a computer system that include one or more physical computer processors and non-transient electronic storage. The computer-implemented method may include: obtaining the target video, extracting one or more frames from the target video, and generating an estimated optical flow based on a displacement of pixels between the one or more frames. The one or more frames may include one or more of a key frame and a target frame.

Background

TECHNICAL FIELD

The present disclosure relates generally to video compression. BRIEF SUMMARY OF THE EMBODIMENTS

Embodiments of the present disclosure include systems and methods of compressing video using machine learning. In accordance with the technology described herein, a computer-implemented method for compressing a target video is disclosed. The computer-implemented method may be implemented in a computer system that may include one or more physical computer processors and non-transient electronic storage. The computer-implemented method may include obtaining, from the non-transient electronic storage, the target video. The computer-implemented method may include extracting, with the one or more physical computer processors, one or more frames from the target video. The one or more frames may include one or more of a key frame and a target frame. The computer-implemented method may also include generating, with the one or more physical computer processors, an estimated optical flow based on a displacement of pixels between the one or more frames.

In embodiments, the displacement of pixels may be between a key frame and/or the target frame.

In embodiments, the computer-implemented method may further include applying, with the one or more physical computer processors, the estimated optical flow to a trained optical flow model to generate a refined optical flow. The trained optical flow model may have been trained by using optical flow training

Claims

1. A computer-implemented method for compressing a target video, the computer-implemented method comprising: determining a first estimated optical flow based on a displacement of pixels between a first reference frame included in the target video and a target frame included in the target video; applying the first estimated optical flow to the first reference frame to produce a first warped target frame; synthesizing, via a first trained machine learning model, an estimate of the target frame based on the first warped target frame; and encoding the target frame based on the estimate of the target frame. || 11. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: determining a first estimated optical flow based on a displacement of pixels between a first reference frame included in a target video and a target frame included in the target video; applying the first estimated optical flow to the first reference frame to produce a first warped target frame; synthesizing, via a first trained machine learning model, an estimate of the target frame based on the first warped target frame; and encoding the target frame based on the estimate of the target frame. || 20. A system, comprising: one or more memories that store instructions, and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to perform the steps of: determining a first estimated optical flow based on a displacement of pixels between a first reference frame included in a target video and a target frame included in the target video; applying the first estimated optical flow to the first reference frame to produce a first warped target frame; synthesizing, via a first trained machine learning model, an estimate of the target frame based on the first warped target frame; and encoding, via a second trained machine learning model, the target frame based on the estimate of the target frame.