Outer Rim Archives
Archives · 2025 · 20250168367

Application (pre-grant publication)

Tunable Hybrid Neural Video Representations

Number
20250168367
Published
2025-05-22
Filed
2024-10-18
Assignee
Disney Enterprises, Inc.
Inventors
Azevedo; Roberto Gerson de Albuquerque et al.
CPC
H04N19/172; H04N19/189; G06N3/045; G06N3/0464; G06N3/084; G06N3/088; G06V10/82; G06V20/46; H04N19/136; H04N19/177
Verdict
Set aside video compression codec research
Source
Google Patents · FreePatentsOnline

Abstract

A system includes a tunable neural network-based video encoder configured to receive a video sequence including multiple video frames, generate a frame-specific embedding of a first video frame of the multiple video frames, and identify one or more group-of-pictures (GOP) features of a subset of the multiple video frames, the subset of the including the first video frame. The tunable neural network-based video encoder is further configured to combine the frame-specific embedding of the first video frame and the one or more GOP features of the first plurality of the plurality of video frames to provide a latent feature corresponding to a compressed version of the first video frame.

Background

BACKGROUND

Video compression is a long-standing and difficult problem that has inspired much research. The main goal of video compression is to represent a digital video (typically a sequence of frames, each being represented by a two-dimensional (2-D) array of pixels, RGB or YUV colors) using the minimum amount of storage while concurrently minimizing loss of quality. Although many advances to traditional video codecs have been made in recent decades, the advent of deep learning has inspired many new approaches that surpass traditional video codecs.

Implicit neural representations (INRs), for example, have attracted significant research interest and have been applied to various domains, including video compression. In addition to exhibiting desirable properties such as fast decoding and the ability of temporal interpolation, INR-based approaches can match or surpass traditional standard video codecs such as Advanced Video Encoding (AVC, also referred to as H.264) and High Efficiency Video Encoding (HEVC) in compression performance. However, these existing approaches to utilizing INRs for video compression only perform well in limited and sometimes highly constrained settings, such as being limited to specific model sizes, fixed aspect ratios, and relatively static video sequences, for instance.

By way of background, early video INRs employed pixel-wise representations that mapped pixel indices to RGB colors but suffered from limited performance and poor

Claims

1. A system comprising: a tunable neural network-based video encoder configured to: receive a video sequence including a plurality of video frames; generate a frame-specific embedding of a first video frame of the plurality of video frames; identify one or more group-of-pictures (GOP) features of a first plurality of the plurality of video frames, the first plurality of the plurality of video frames including the first video frame; and combine the frame-specific embedding of the first video frame and the one or more GOP features of the first plurality of the plurality of video frames to provide a latent feature corresponding to a compressed version of the first video frame. || 9. A system comprising: a neural network-based video decoder configured to: receive a latent feature corresponding to a compressed version of a first video frame of a first plurality of a plurality of video frames included in a video sequence, the latent feature being a combination of a frame-specific embedding of the first video frame with one or more group-of-pictures (GOP) features of the first plurality of the plurality of video frames; and decode the latent feature to provide an uncompressed video frame corresponding to the first video frame. || 13. A method for execution by a tunable neural network-based video encoder, the method comprising: receiving a video sequence including a plurality of video frames; generating a frame-specific embedding of a first video frame of the plurality of video frames; identifying one or more group-of-pictures (GOP) features of a first plurality of the plurality of video frames, the first plurality of the plurality of video frames including the first video frame; and combining the frame-specific embedding of the first video frame and the one or more GOP features of the first plurality of the plurality of video frames to provide a latent feature corresponding to a compressed version of the first video frame.