Outer Rim Archives
Archives · 2024 · 20240305801

Application (pre-grant publication)

CONTEXTUAL VIDEO COMPRESSION FRAMEWORK WITH SPATIAL-TEMPORAL CROSS-COVARIANCE TRANSFORMERS

Number
20240305801
Published
2024-09-12
Filed
2023-07-07
Assignee
Disney Enterprises, Inc.
Inventors
Chen; Zhenghao et al.
CPC
H04N19/172; H04N19/42; H04N19/513; H04N19/593; H04N19/91
Verdict
Set aside video compression codec, plumbing
Source
Google Patents · FreePatentsOnline

Abstract

In some embodiments, a system includes a first component to extract temporal features from a current frame being coded and a previous frame of a video. A second component uses a first transformer to fuse spatial features from the current frame with the temporal features to generate spatio-temporal features as first output. A third component uses a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output. A fourth component uses a third transformer to reconstruct the current frame based on the first output that is processed using the second output and the temporal features.

Background

BACKGROUND

Video compression reduces the amount of data that is stored or transmitted for videos. Achieving an efficient reduction in data is important considering the increasing demand for storing and transmitting videos. Video compression may attempt to exploit spatial redundancy between pixels in the same video frame or temporal redundancy between pixels in multiple video frames. Some video compression methods may focus on improving either temporal information or spatial information separately. Then, these methods may combine spatial information and temporal information using simple operations, such as concatenation or subtraction. However, these operations may only partially exploit the spatial-temporal redundancies.

Claims

1. A system comprising: a first component to extract temporal features from a current frame being coded and a previous frame of a video; a second component that uses a first transformer to fuse spatial features from the current frame with the temporal features to generate spatio-temporal features as first output; a third component that uses a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output; and a fourth component that uses a third transformer to reconstruct the current frame based on the first output that is processed using the second output and the temporal features. || 15. A method comprising: extracting temporal features from a current frame being coded and a previous frame of a video; using a first transformer to fuse spatial features from the current frame with the temporal features to generate spatio-temporal features as first output; using a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output; and using a third transformer to reconstruct the current frame based on the first output that is processed using the second output and the temporal features. || 20. An apparatus comprising: one or more computer processors; and a computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable for: extracting temporal features from a current frame being coded and a previous frame of a video; using a first transformer to fuse spatial features from the current frame with the temporal features to generate spatio-temporal features as first output; using a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output; and using a third transformer to reconstruct the current frame based on the first output that is processed using the second output and the temporal features.