Application (pre-grant publication)
CONTEXTUAL VIDEO COMPRESSION FRAMEWORK WITH SPATIAL-TEMPORAL CROSS-COVARIANCE TRANSFORMERS
- Number
- 20240305801
- Published
- 2024-09-12
- Filed
- 2023-07-07
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Chen; Zhenghao et al.
- CPC
- H04N19/172; H04N19/42; H04N19/513; H04N19/593; H04N19/91
- Verdict
- Set aside video compression codec, plumbing
- Source
- Google Patents · FreePatentsOnline
Abstract
In some embodiments, a system includes a first component to extract temporal features from a current frame being coded and a previous frame of a video. A second component uses a first transformer to fuse spatial features from the current frame with the temporal features to generate spatio-temporal features as first output. A third component uses a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output. A fourth component uses a third transformer to reconstruct the current frame based on the first output that is processed using the second output and the temporal features.
Background
BACKGROUND
Video compression reduces the amount of data that is stored or transmitted for videos. Achieving an efficient reduction in data is important considering the increasing demand for storing and transmitting videos. Video compression may attempt to exploit spatial redundancy between pixels in the same video frame or temporal redundancy between pixels in multiple video frames. Some video compression methods may focus on improving either temporal information or spatial information separately. Then, these methods may combine spatial information and temporal information using simple operations, such as concatenation or subtraction. However, these operations may only partially exploit the spatial-temporal redundancies.