Outer Rim Archives
Archives · 2023 · 11765360

Granted patent

Codec rate distortion compensating downsampler

Number
11765360
Published
2023-09-19
Filed
2021-10-13
Assignee
Disney Enterprises, Inc.
Inventors
Schroers; Christopher Richard et al.
CPC
G06N3/0464; G06N3/08; G06T3/4046; G06T9/002; H04N19/132; H04N19/147; H04N19/149; H04N19/154; H04N19/184
Verdict
Set aside video codec compression, plumbing
Source
Google Patents · FreePatentsOnline

Abstract

A system includes a machine learning (ML) model-based video downsampler configured to receive an input video sequence having a first display resolution, and to map the input video sequence to a lower resolution video sequence having a second display resolution lower than the first display resolution. The system also includes a neural network-based (NN-based) proxy video codec configured to transform the lower resolution video sequence into a decoded proxy bitstream. In addition, the system includes an upsampler configured to produce an output video sequence using the decoded proxy bitstream.

Background

BACKGROUND (1) Downsampling is an operation in content streaming systems to produce different representations in terms of bit rate and resolution available to different types of client devices. In modern streaming systems, the streaming server provides different encoding representations in terms of resolutions and bitrates, so that the client device can dynamically download the representation that best matches its playback context (e.g., display size and network conditions). In order to provide such representations, the streaming server needs to downsample the source video to different resolutions before encoding. That downsampling may be performed with filters that are not perceptually optimal.

Claims

1. A system comprising: (a) a machine learning (ML) model-based video downsampler configured to: receive an input video sequence having a first display resolution; and map the input video sequence to a lower resolution video sequence having a second display resolution lower than the first display resolution; (b) a neural network-based (NN-based) proxy video codec configured to transform the lower resolution video sequence into a decoded proxy bitstream, wherein the NN-based proxy video codec is pre-trained to replicate a rate distortion characteristic of a standard video codec; and (c) an upsampler configured to produce an output video sequence using the decoded proxy bitstream. || 6. A system of comprising: (a) a machine learning (ML) model-based video downsampler configured to: receive an input video sequence having a first display resolution; and map the input video sequence to a lower resolution video sequence having a second display resolution lower than the first display resolution; (b) a neural network-based (NN-based) proxy video codec configured to transform the lower resolution video sequence into a decoded proxy bitstream; and (c) an upsampler configured to produce an output video sequence using the decoded proxy bitstream wherein the ML model-based video downsampler is trained using the input video sequence, the output video sequence, and an objective function based on an estimated rate of the lower resolution video sequence and a plurality of perceptual loss functions. || 10. A method for training a machine learning (ML) model-based video downsampler, the method comprising: providing, to the ML model-based video downsampler, an input video sequence having a first display resolution; mapping, using the ML model-based video downsampler, the input video sequence to a lower resolution video sequence having a second display resolution lower than the first display resolution; transforming, using a neural network-based (NN-based) proxy video codec, the lower resolution video sequence into a decoded proxy bitstream; producing, using an upsampler receiving the decoded proxy bitstream, an output video sequence corresponding to the input video sequence and having a display resolution higher than the second display resolution; and training the ML model-based video downsampler using the input video sequence, the output video sequence, and an objective function based on an estimated rate of the lower resolution video sequence and a plurality of perceptual loss functions. || 17. A video processing system comprising: a processing hardware and a system memory storing a video codec and a trained ML model-based video downsampler that has been trained using a neural network-based (NN-based) proxy video codec configured to replicate a rate distortion characteristic of the video codec; the processing hardware configured to: receive an input video sequence having a first display resolution; map, using the trained ML model-based video downsampler, the input video sequence to a lower resolution video sequence having a second display resolution lower than the first display resolution; transform, using the video codec, the lower resolution video sequence into a decoded bitstream; and output the decoded bitstream. || 20. A video processing system comprising: a simulation module including a neural network-based (NN-based) proxy video codec; and a processing hardware and a system memory storing a video codec and a trained ML model-based video downsampler that has been trained using the NN-based proxy video codec; the processing hardware configured to: receive an input video sequence having a first display resolution; map, using the trained ML model-based video downsampler, the input video sequence to a lower resolution video sequence having a second display resolution lower than the first display resolution; transform, using the video codec, the lower resolution video sequence into a decoded bitstream; and output the decoded bitstream.