Outer Rim Archives
Archives · 2026 · 12634488

Granted patent

Tunable hybrid neural video representations

Number
12634488
Published
2026-05-19
Filed
2024-10-18
Assignee
Disney Enterprises, Inc.
Inventors
Azevedo; Roberto Gerson de Albuquerque, Schroers; Christopher Richard, Labrozzi; Scott, Saethre; Jens Eirik, Xue; Yuanyì
CPC
H04N19/172; H04N19/189; G06N3/045; G06N3/0464; G06N3/084; G06N3/088; G06V10/82; G06V20/46; H04N19/136; H04N19/177
Verdict
Set aside codec
Source
Google Patents · FreePatentsOnline

The keeper's note

A system includes a tunable neural network-based video encoder configured to receive a video sequence including multiple video frames, generate a frame-specific embedding of a first video frame of the multiple video fra…

Abstract

A system includes a tunable neural network-based video encoder configured to receive a video sequence including multiple video frames, generate a frame-specific embedding of a first video frame of the multiple video frames, and identify one or more group-of-pictures (GOP) features of a subset of the multiple video frames, the subset of the including the first video frame. The tunable neural network-based video encoder is further configured to combine the frame-specific embedding of the first video frame and the one or more GOP features of the first plurality of the plurality of video frames to provide a latent feature corresponding to a compressed version of the first video frame.

Background

BACKGROUND (1) Video compression is a long-standing and difficult problem that has inspired much research. The main goal of video compression is to represent a digital video (typically a sequence of frames, each being represented by a two-dimensional (2-D) array of pixels, RGB or YUV colors) using the minimum amount of storage while concurrently minimizing loss of quality. Although many advances to traditional video codecs have been made in recent decades, the advent of deep learning has inspired many new approaches that surpass traditional video codecs. (2) Implicit neural representations (INRs), for example, have attracted significant research interest and have been applied to various domains, including video compression. In addition to exhibiting desirable properties such as fast decoding and the ability of temporal interpolation, INR-based approaches can match or surpass traditional standard video codecs such as Advanced Video Encoding (AVC, also referred to as H.264) and High Efficiency Video Encoding (HEVC) in compression performance. However, these existing approaches to utilizing INRs for video compression only perform well in limited and sometimes highly constrained settings, such as being limited to specific model sizes, fixed aspect ratios, and relatively static video sequences, for instance. (3) By way of background, early video INRs employed pixel-wise representations that mapped pixel indices to RGB colors but suffered from limited performance and poor decoding

Claims

1. A system comprising: a tunable neural network-based video encoder configured to: receive a video content including a plurality of video frames; generate a frame-specific embedding of a first video frame of the plurality of video frames; identify one or more group-of-pictures (GOP) features of a first plurality of the plurality of video frames, the first plurality of the plurality of video frames including the first video frame; determine a first weight and a second weight for applying to the frame-specific embedding and the one or more GOP features, respectively, as levers for content-specific fine-tuning; apply a first weight to the frame-specific embedding to obtain a weighted frame-specific embedding; apply a second weight to the one or more GOP features to obtain weighted one or more GOP features; and combine the weighted frame-specific embedding of the first video frame and the weighted one or more GOP features of the first plurality of the plurality of video frames to provide a latent feature corresponding to a compressed version of the first video frame. || 7. A system comprising: a neural network-based video decoder configured to: receive a latent feature corresponding to a compressed version of a first video frame of a first plurality of a plurality of video frames included in a video content, the latent feature being a combination of a weighted frame-specific embedding of the first video frame with weighted one or more group-of-pictures (GOP) features of the first plurality of the plurality of video frames, wherein the weighted frame-specific embedding is a product of applying a first weight to a frame-specific embedding and the weighted one or more GOP features are products of applying a second weight to the one or more GOP features, wherein the first weight and the second weight are selected as levers for content-specific fine-tuning; and decode the latent feature to provide an uncompressed video frame corresponding to the first video frame. || 10. A method for execution by a tunable neural network-based video encoder, the method comprising: receiving a video content including a plurality of video frames; generating a frame-specific embedding of a first video frame of the plurality of video frames; identifying one or more group-of-pictures (GOP) features of a first plurality of the plurality of video frames, the first plurality of the plurality of video frames including the first video frame; determining a first weight and a second weight for applying to the frame-specific embedding and the one or more GOP features, respectively, as levers for content-specific fine-tuning; applying a first weight to the frame-specific embedding to obtain a weighted frame-specific embedding; applying a second weight to the one or more GOP features to obtain a weighted one or more GOP features; and combining the weighted frame-specific embedding of the first video frame and the weighted one or more GOP features of the first plurality of the plurality of video frames to provide a latent feature corresponding to a compressed version of the first video frame.