Outer Rim Archives
Archives · 2023 · 11544606

Granted patent

Machine learning based video compression

Number
11544606
Published
2023-01-03
Filed
2019-01-22
Assignee
Disney Enterprises, Inc.
Inventors
Mandt; Stephan Marcel et al.
CPC
G06N3/088; G06N3/0464; G06N3/0495; H04N19/124; G06N3/0475; H04N19/503; G06N20/00; G06N3/047; G06N3/0442; G06N3/044; G06N3/04; H04N19/91; H04N19/42; G06N7/01; G06N3/08; G06N3/0455; G06T9/002; H04N19/46; G06N3/045
Verdict
Set aside video compression codec, plumbing
Source
Google Patents · FreePatentsOnline

Abstract

Systems and methods for compressing target content are disclosed. In one embodiment, a system may include non-transient electronic storage and one or more physical computer processors. The one or more physical computer processors may be configured by machine-readable instructions to obtain the target content comprising one or more frames, wherein a given frame comprises one or more features. The one or more physical computer processors may be configured by machine-readable instructions to obtain a conditioned network. The one or more physical computer processors may be configured by machine-readable instructions to generate decoded target content by applying the conditioned network to the target content.

Background

TECHNICAL FIELD (1) The present disclosure relates generally to video compression. BRIEF SUMMARY (2) Embodiment of the disclosure are directed to systems and methods for compressing video. (3) In one embodiment, a system may be configured for compressing content. The system may include non-transient electronic storage and one or more physical computer processors configured by machine-readable instructions to perform a number of operations. One operation may be to obtain, from the non-transient electronic storage, the target content comprising one or more frames. A given frame may include one or more features. Another operation may be to obtain, from the non-transient electronic storage, a conditioned network. The conditioned network may have been trained by training an initial network using training content. The conditioned network may include one or more encoders, one or more quantizers, and one or more decoders. The training content may include one or more training frames. A given training frame may include one or more training features. Another such operation may be to generate, with the one or more physical computer processors, decoded target content by applying the conditioned network to the target content. The conditioned network may generate a latent space of the target content. The target content may include one or more local variables and one or more global variables. (4) In embodiments, applying the conditioned network may include encoding, with the one or more phys

Claims

1. A system configured for compressing target content, the system comprising: non-transient electronic storage; one or more physical computer processors configured by machine-readable instructions to: obtain, from the non-transient electronic storage, the target content comprising one or more frames, wherein a given frame comprises one or more features; obtain, from the non-transient electronic storage, a conditioned network, the conditioned network having been trained by training an initial network using training content, wherein the conditioned network comprises one or more encoders, one or more quantizers, and one or more decoders, and wherein the training content comprises one or more training frames, and wherein a given training frame comprises one or more training features; apply, with the one or more physical computer processors, the conditioned network to the target content to generate a latent space of the target content comprising one or more local variables, one or more global variables, and a plurality of distributions corresponding to the latent space; and quantize the one or more local variables and the one or more global variables based on the plurality of distributions corresponding to the latent space to generate encoded target content. || 8. A computer-implemented method for training an initial network to simultaneously learn how to refine a latent space using training content and how to refine a plurality of distributions of the latent space using the training content, the method being implemented in a computer system that comprises non-transient electronic storage and one or more physical computer processors, comprising: obtaining, from the non-transient electronic storage, training content comprising one or more training frames, wherein a given training frame comprises one or more training features; obtaining, from the non-transient electronic storage, the initial network, the initial network comprising one or more encoders, one or more quantizers, and one or more decoders; and generating, with the one or more physical computer processors, a conditioned network by training the initial network using the training content, the conditioned network comprising the one or more encoders, the one or more quantizers, and the one or more decoders, wherein the conditioned network is trained to receive target content and generate encoded target content comprising a quantized latent space that includes one or more quantized local variables and one or more quantized global variables. || 14. A computer-implemented method for compressing target content, the method being implemented in a computer system that comprises non-transient electronic storage and one or more physical computer processors, comprising: obtaining, from the non-transient electronic storage, the target content comprising one or more frames, wherein a given frame comprises one or more features; encoding, with the one or more physical computer processors, the target content to generate one or more local variables and one or more global variables; and generating, with the one or more physical computer processors, a latent space, the latent space comprising the one or more local variables and the one or more global variables, wherein the one or more local variables are based on the one or more features in the given frame, and wherein the one or more global variables are based on one or more features common to a plurality of frames of the target content; generating, with the one or more physical computer processors, a plurality of distributions corresponding to the latent space; and quantizing the one or more local variables and the one or more global variables based on the plurality of distributions corresponding to the latent space.