Outer Rim Archives
Archives · 2025 · 20250211758

Application (pre-grant publication)

Codec Rate Distortion Compensating Downsampler

Number
20250211758
Published
2025-06-26
Filed
2025-03-12
Assignee
Disney Enterprises, Inc.
Inventors
Schroers; Christopher Richard et al.
CPC
G06N3/0464; H04N19/184; H04N19/154; H04N19/149; H04N19/132; G06T9/002; H04N19/147; G06T3/4046; G06N3/08
Verdict
Set aside codec rate-distortion, plumbing
Source
Google Patents · FreePatentsOnline

Abstract

A system includes a machine learning (ML) model-based video downsampler configured to receive an input video sequence having a first display resolution, and to map the input video sequence to a lower resolution video sequence having a second display resolution lower than the first display resolution. The system also includes a neural network-based (NN-based) proxy video codec configured to transform the lower resolution video sequence into a decoded proxy bitstream. In addition, the system includes an upsampler configured to produce an output video sequence using the decoded proxy bitstream.

Background

BACKGROUND

Downsampling is an operation in content streaming systems to produce different representations in terms of bit rate and resolution available to different types of client devices. In modern streaming systems, the streaming server provides different encoding representations in terms of resolutions and bitrates, so that the client device can dynamically download the representation that best matches its playback context (e.g., display size and network conditions). In order to provide such representations, the streaming server needs to downsample the source video to different resolutions before encoding. That downsampling may be performed with filters that are not perceptually optimal.

Claims

23. A video processing system comprising: an upsampler; a machine learning (ML) model-based video downsampler trained using a plurality of perceptual loss functions; and a processing hardware configured to: receive an input video sequence having a first display resolution; extract a content sample of the input video sequence; map, using the trained ML model-based video downsampler, the content sample to a lower resolution sample; transform, using one of a video codec or a neural network-based (NN-based) proxy video codec, the lower resolution sample into a decoded sample bitstream; predict, using the upsampler and the decoded sample bitstream, an output sample corresponding to the content sample; and modify, based on the predicted output sample, one or more parameters of the trained ML model-based video downsampler. || 31. A method for use by a video processing system including an upsampler, and a machine learning (ML) model-based video downsampler trained using a plurality of perceptual loss functions, the method comprising: receiving an input video sequence having a first display resolution; extracting a content sample of the input video sequence; mapping, using the trained ML model-based video downsampler, the content sample to a lower resolution sample; transforming, using one of a video codec or a neural network-based (NN-based) proxy video codec, the lower resolution sample into a decoded sample bitstream; predicting, using the upsampler and the decoded sample bitstream, an output sample corresponding to the content sample; and modifying, based on the predicted output sample, one or more parameters of the trained ML model-based video downsampler.