Outer Rim Archives
Archives · 2024 · 20240283957

Application (pre-grant publication)

Microdosing For Low Bitrate Video Compression

Number
20240283957
Published
2024-08-22
Filed
2024-05-02
Assignee
Disney Enterprises, Inc.
Inventors
Djelouah; Abdelaziz et al.
CPC
G06N3/045; G06N3/0455; G06N3/0464; G06N3/047; G06N3/0475; G06N3/0495; G06N3/088; G06N3/09; G06N3/094; G06N3/096; H04N19/42; H04N19/44
Verdict
Set aside video compression codec, plumbing
Source
Google Patents · FreePatentsOnline

Abstract

A system includes a machine learning (ML) model-based video encoder configured to receive an uncompressed video sequence including multiple video frames, determine, from among the multiple video frames, a first video frame subset and a second video frame subset, encode the first video frame subset to produce a first compressed video frame subset, and identify a first decompression data for the first compressed video frame subset. The ML model-based video encoder is further configured to encode the second video frame subset to produce a second compressed video frame subset, and identify a second decompression data for the second compressed video frame subset. The first decompression data is specific to decoding the first compressed video frame subset but not the second compressed video frame subset, and the second decompression data is specific to decoding the second compressed video frame subset but not the first compressed video frame subset.

Background

BACKGROUND

Video content represents the majority of total Internet traffic and is expected to increase even more as spatial resolution frame rate, and color depth of videos increase and more users adopt streaming services. Although existing codecs have achieved impressive performance, they have been engineered to the point where adding further small improvements is unlikely to meet future demands. Consequently, exploring fundamentally different ways to perform video coding may advantageously lead to a new class of video codecs with improved performance and flexibility.

For example, one advantage of using a trained machine learning (ML) model, such as a neural network (NN), in the form of a generative adversarial network (GAN) for example, to perform video compression is that it enables the ML model to infer visual details that it would otherwise be costly in terms of data transmission, to obtain. However, the model size remains an important issue in current state-of-the-art proposals and existing solutions require significant computation effort on the decoding side. That is to say, one significant drawback of existing GAN-based compression frameworks is that they typically require large decoder models that are sometimes trained on private datasets. Therefore, retraining these models to their original performance is not generally possible, and even when the training data is available, retraining the model would be complicated and time consuming. Moreover, the mem

Claims

21: A method for use by a machine learning (ML) model-based video decoder, the method comprising: receiving, by a degradation-aware block based Micro-Residual-Network (MicroRN) defined by a number of hidden channels and a number of degradation-aware blocks of the MicroRN, a first compressed video frame subset; receiving, by the MicroRN, first decompression data for the first compressed video frame subset; decoding, by the MicroRN, the first compressed video frame subset using the first decompression data; receiving, by the MicroRN, a second compressed video frame subset; receiving, by the MicroRN, second decompression data for the second compressed video frame subset; and decoding, by the MicroRN, the second compressed video frame subset using the second decompression data, without utilizing a residual network of a generative adversarial network (GAN) trained decoder. || 26: A machine learning (ML) model-based video decoder comprising: a degradation-aware block based Micro-Residual-Network (MicroRN) defined by a number of hidden channels and a number of degradation-aware blocks of the MicroRN; the MicroRN being configured to: receive a first compressed video frame subset; receive first decompression data for the first compressed video frame subset; decode the first compressed video frame subset using the first decompression data; receive a second compressed video frame subset; receive second decompression data for the second compressed video frame subset; and decode the second compressed video frame subset using the second decompression data, without utilizing a residual network of a generative adversarial network (GAN) trained decoder. || 33: A method for use by a machine learning (ML) model-based video decoder in a system including an ML model-based video encoder configured to encode a first video frame subset to produce a first compressed video frame subset, identify first decompression data for the first compressed video frame subset, encode a second video frame subset to produce a second compressed video frame subset, and identify second decompression data for the second compressed video frame subset, the method comprising: receiving, by a degradation-aware block based Micro-Residual-Network (MicroRN) defined by a number of hidden channels and a number of degradation-aware blocks of the MicroRN, the first compressed video frame subset; receiving, by the MicroRN, the first decompression data; decoding, by the MicroRN, the first compressed video frame subset using the first decompression data; and receiving, by the MicroRN, the second compressed video frame subset; receiving, by the MicroRN, the second decompression data; and decoding, by the MicroRN, the second compressed video frame subset using the second decompression data, without utilizing a residual network of a generative adversarial network (GAN) trained decoder.