Outer Rim Archives
Archives · 2024 · 12120359

Granted patent

Machine learning model-based video compression

Number
12120359
Published
2024-10-15
Filed
2022-03-25
Assignee
Disney Enterprises, Inc.
Inventors
Djelouah; Abdelaziz et al.
CPC
H04N19/503; H04N19/89; G06N3/045; G06N3/0455; G06N3/0464; G06N3/0475; G06N3/088; G06N3/09; G06N3/094
Verdict
Set aside video compression codec, plumbing
Source
Google Patents · FreePatentsOnline

Abstract

A system processing hardware executes a machine learning (ML) model-based video compression encoder to receive uncompressed video content and corresponding motion compensated video content, compare the uncompressed and motion compensated video content to identify an image space residual, transform the image space residual to a latent space representation of the uncompressed video content, and transform, using a trained image compression ML model, the motion compensated video content to a latent space representation of the motion compensated video content. The ML model-based video compression encoder further encodes the latent space representation of the image space residual to produce an encoded latent residual, encodes, using the trained image compression ML model, the latent space representation of the motion compensated video content to produce an encoded latent video content, and generates, using the encoded latent residual and the encoded latent video content, a compressed video content corresponding to the uncompressed video content.

Background

BACKGROUND (1) Video content represents the majority of total Internet traffic and is expected to increase even more as spatial resolution frame rate, and color depth of videos increase and more users adopt earning services. Although existing codecs have achieved impressive performance, they have been engineered to the point where adding further small improvements is unlikely to meet future demands. Consequently, exploring fundamentally different ways to perform video coding may advantageously lead to a new class of video codecs with improved performance and flexibility. (2) For example, one advantage of using a trained machine learning (ML) model, such as a neural network (NN), in the form of a generative adversarial network (GAN) for example, to perform video compression is that it enables the ML model to infer visual details that it would otherwise be costly in terms of data transmission, obtain. However, training ML models such as GANs is typically challenging because the training alternates between minimization and maximization steps to converge to a saddle point of the loss function. The task becomes more challenging when considering the temporal domain and the increased complexity it introduces.

Claims

1. A system comprising: a computing platform including a processing hardware and a system memory storing a machine learning (ML) model-based video compression encoder and a trained image compression ML model; the processing hardware configured to execute the ML model-based video compression encoder to: receive an uncompressed video content and a motion compensated video content corresponding to the uncompressed video content; compare the uncompressed video content with the motion compensated video content to identify an image space residual corresponding to the uncompressed video content; transform the image space residual to a latent space representation of the image space residual; receive, using the trained image compression ML model, the motion compensated video content; transform, using the trained image compression ML model, the motion compensated video content to a latent space representation of the motion compensated video content; encode the latent space representation of the image space residual to produce an encoded latent residual; encode, using the trained image compression ML model, the latent space representation of the motion compensated video content to produce an encoded latent video content; and generate, using the encoded latent residual and the encoded latent video content, a compressed video content corresponding to the uncompressed video content. || 7. A method for use by a system including a computing platform having a processing hardware and a system memory storing a machine learning (ML) model-based video compression encoder and a trained image compression ML model, the method comprising: receiving, by the ML model-based video compression encoder executed by the processing hardware, an uncompressed video content and a motion compensated video content corresponding to the uncompressed video content; comparing, by the ML model-based video compression encoder executed by the processing hardware, the uncompressed video content with the motion compensated video content, thereby identifying an image space residual corresponding to the uncompressed video content; transforming, by the ML model-based video compression encoder executed by the processing hardware, the image space residual to a latent space representation of the image space residual; receiving, by the trained image compression ML model executed by the processing hardware, the motion compensated video content; transforming, by the trained image compression ML model executed by the processing hardware, the motion compensated video content to a latent space representation of the motion compensated video content; encoding, by the ML model-based video compression encoder executed by the processing hardware, the latent space representation of the image space residual to produce an encoded latent residual; encoding, by the trained image compression ML model executed by the processing hardware, the latent space representation of the motion compensated video content to produce an encoded video content; and generating, by the ML model-based video compression encoder executed by the processing hardware and using the encoded latent residual and the encoded latent video content, a compressed video content corresponding to the uncompressed video content. || 13. A system comprising: a computing platform including a processing hardware and a system memory storing a machine learning (ML) model-based video compression encoder and a trained image compression ML model; the processing hardware configured to execute the ML model-based video compression encoder to: receive, using the trained image compression ML model, an uncompressed video content and a motion compensated video content corresponding to the uncompressed video content; transform, using the trained image compression ML model, the uncompressed video content to a first latent space representation of the uncompressed video content; transform, using the trained image compression ML model, the motion compensated video content to a second latent space representation of the uncompressed video content; and generate a bitstream for transmitting a compressed video content corresponding to the uncompressed video content based on the first latent space representation and the second latent space representation.