Outer Rim Archives
Archives · 2025 · 12278969

Granted patent

Codec rate distortion compensating downsampler

Number
12278969
Published
2025-04-15
Filed
2023-08-04
Assignee
Disney Enterprises, Inc.
Inventors
Schroers; Christopher Richard et al.
CPC
G06N3/0464; H04N19/132; H04N19/147; G06N3/08; H04N19/154; H04N19/184; G06T9/002; G06T3/4046; H04N19/149
Verdict
Set aside codec rate-distortion, plumbing
Source
Google Patents · FreePatentsOnline

Abstract

A system includes a machine learning (ML) model-based video downsampler configured to receive an input video sequence having a first display resolution, and to map the input video sequence to a lower resolution video sequence having a second display resolution lower than the first display resolution. The system also includes a neural network-based (NN-based) proxy video codec configured to transform the lower resolution video sequence into a decoded proxy bitstream. In addition, the system includes an upsampler configured to produce an output video sequence using the decoded proxy bitstream.

Background

BACKGROUND (1) Downsampling is an operation in content streaming systems to produce different representations in terms of bit rate and resolution available to different types of client devices. In modern streaming systems, the streaming server provides different encoding representations in terms of resolutions and bitrates, so that the client device can dynamically download the representation that best matches its playback context (e.g., display size and network conditions). In order to provide such representations, the streaming server needs to downsample the source video to different resolutions before encoding. That downsampling may be performed with filters that are not perceptually optimal.

Claims

1. A video processing system comprising: an upsampler; a video codec; a trained machine learning (ML) model-based video downsampler trained using a neural network-based (NN-based) proxy video codec; and a processing hardware configured to: receive an input video sequence having a first display resolution; extract a content sample of the input video sequence; map, using the trained ML model-based video downsampler, the content sample to a lower resolution sample; transform, using one of the video codec or the NN-based proxy video codec, the lower resolution sample into a decoded sample bitstream; predict, using the upsampler and the decoded sample bitstream, an output sample corresponding to the content sample; and modify, based on the predicted output sample, one or more parameters of the trained ML model-based video downsampler; wherein the ML model-based video downsampler is trained using the input video sequence, the output sample, and an objective function based on an estimated rate of the lower resolution sample and a plurality of perceptual loss functions. || 10. A method for use by a video processing system including an upsampler, a video codec, and a trained machine learning (ML) model-based video downsampler trained using a neural network-based (NN-based) proxy video codec, the method comprising: receiving an input video sequence having a first display resolution; extracting a content sample of the input video sequence; mapping, using the trained ML model-based video downsampler, the content sample to a lower resolution sample; transforming, using one of the video codec or the NN-based proxy video codec, the lower resolution sample into a decoded sample bitstream; predicting, using the upsampler and the decoded sample bitstream, an output sample corresponding to the content sample; and modifying, based on the predicted output sample, one or more parameters of the trained ML model-based video downsampler; wherein the ML model-based video downsampler is trained using the input video sequence, the output sample, and an objective function based on an estimated rate of the lower resolution sample and a plurality of perceptual loss functions.