In some embodiments, a method receives a first video. The first video includes frames that were generated using frame interpolation. A feature extractor extracts first features from frames of the first video. The first features are extracted from a plurality of levels of a network of the feature extractor. A spatio-temporal processing system analyzes the first features spatially and temporally to determine spatial and temporal features for the plurality of levels. The method combines the spatial and temporal features from the plurality of levels to determine a score that measures a quality of the first video.
BACKGROUND
Video frame interpolation generates new frames from existing frames of a video. The new frames may be used to up sample the video frame rate. Video frame interpolation can introduce artifacts that degrade the perceived quality of the video. A system can attempt to assess the quality of the video using metrics that compare pixel values or low level visual patterns to perform the quality assessment, such as peak signal-to-noise ratio (PSNR) or structural similarity index measure (SSIM). However, assessing the quality of the video using these metrics may not properly evaluate the artifacts introduced by video frame interpolation. For example, the metrics are not specifically designed to recognize and evaluate the video frame interpolation artifacts, which may include temporal artifacts.
1. A method comprising: receiving a first video, wherein the first video includes frames that were generated using frame interpolation; extracting, using a feature extractor, first features from frames of the first video, wherein first features are extracted from a plurality of levels of a network of the feature extractor; analyzing, using a spatio-temporal processing system, the first features spatially and temporally to determine spatial and temporal features for the plurality of levels; and combining the spatial and temporal features from the plurality of levels to determine a score that measures a quality of the first video. ||
15. A non-transitory computer-readable storage medium having stored thereon computer executable instructions, which when executed by a computing device, cause the computing device to be operable for: receiving a first video, wherein the first video includes frames that were generated using frame interpolation; extracting, using a feature extractor, first features from frames of the first video, wherein first features are extracted from a plurality of levels of a network of the feature extractor; analyzing, using a spatio-temporal processing system, the first features spatially and temporally to determine spatial and temporal features for the plurality of levels; and combining the spatial and temporal features from the plurality of levels to determine a score that measures a quality of the first video. ||
16. A method comprising: receiving a first video, wherein the first video includes frames that were generated using frame interpolation; receiving a second video, wherein the second video does not include frames that were generated using frame interpolation; extracting, using a feature extractor, first features from frames of the first video and second features from frames of the second video, wherein the first features and the second features are extracted from the plurality of levels of the network of the feature extractor; combining the second features with the first features to determine concatenated features; analyzing the concatenated features spatially and temporally to determine spatial and temporal concatenated features for the plurality of levels; and combining the spatial and temporal concatenated features from the plurality of levels to determine a score that measures a quality of the first video.