In some embodiments, a system includes a first component to extract temporal features from a current frame being coded and a previous frame of a video. A second component uses a first transformer to fuse spatial features from the current frame with the temporal features to generate spatio-temporal features as first output. A third component uses a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output. A fourth component uses a third transformer to reconstruct the current frame based on the first output that is processed using the second output and the temporal features.
BACKGROUND (1) Video compression reduces the amount of data that is stored or transmitted for videos. Achieving an efficient reduction in data is important considering the increasing demand for storing and transmitting videos. Video compression may attempt to exploit spatial redundancy between pixels in the same video frame or temporal redundancy between pixels in multiple video frames. Some video compression methods may focus on improving either temporal information or spatial information separately. Then, these methods may combine spatial information and temporal information using simple operations, such as concatenation or subtraction. However, these operations may only partially exploit the spatial-temporal redundancies.
1. A system comprising: a first component to extract temporal features from a current frame being coded and a previous frame of a video, wherein three-dimensional based joint features are determined using the temporal features and spatial features from the current frame; a second component that uses a first transformer to receive the three-dimensional based joint features as input and fuse the spatial features from the current frame with the temporal features to generate spatio-temporal features as first output; a third component that uses a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output, wherein the second transformer is used to fuse the spatio-temporal features with the at least a portion of the temporal features to output fused spatio-temporal features that are entropy encoded to generate the second output; and a fourth component that uses a third transformer to reconstruct the current frame, wherein the first output is processed using the second output to generate third output, and wherein the third transformer fuses the temporal features with the third output. ||
15. A method comprising: extracting temporal features from a current frame being coded and a previous frame of a video, wherein three-dimensional based joint features are determined using the temporal features and spatial features from the current frame; using a first transformer to receive the three-dimensional based joint features as input and fuse the spatial features from the current frame with the temporal features to generate spatio-temporal features as first output; using a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output, wherein the second transformer is used to fuse the spatio-temporal features with the at least a portion of the temporal features to output fused spatio-temporal features that are entropy encoded to generate the second output; and using a third transformer to reconstruct the current frame, wherein the first output is processed using the second output to generate a third output, and wherein the third transformer fuses the temporal features with the third output. ||
20. An apparatus comprising: one or more computer processors; and a computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable for: extracting temporal features from a current frame being coded and a previous frame of a video, wherein three-dimensional based joint features are determined using the temporal features and spatial features from the current frame; using a first transformer to receive the three-dimensional based joint features as input and fuse the spatial features from the current frame with the temporal features to generate spatio-temporal features as first output; using a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output, wherein the second transformer is used to fuse the spatio-temporal features with the at least a portion of the temporal features to output fused spatio-temporal features that are entropy encoded to generate the second output; and using a third transformer to reconstruct the current frame, wherein the first output is processed using the second output to generate a third output and wherein the third transformer fuses the temporal features with the third output.