Outer Rim Archives
Archives · 2026 · 20260179202

Application (pre-grant publication)

COMBINING STATE SPACE MODELS AND CONVOLUTIONAL NEURAL NETWORKS FOR GENERIC VIDEO QUALITY ASSESSMENT

Number
20260179202
Published
2026-06-25
Filed
2025-11-26
Assignee
Disney Enterprises, Inc.
Inventors
ZHANG; Yang, YANG; Felix, SCHROERS; Christopher Richard
CPC
G06T7/0002; G06T2207/10016; G06T2207/20021; G06T2207/20081; G06T2207/20084; G06T2207/30168
Verdict
Set aside streaming, codec, ad-tech
Source
Google Patents · FreePatentsOnline

The keeper's note

Systems and methods are disclosed for automated video quality assessment using a hybrid neural network architecture.

Abstract

Systems and methods are disclosed for automated video quality assessment using a hybrid neural network architecture. A sequence of video frames is received and partitioned into fragments, which are further subdivided into patches and encoded as tokens. The tokens are processed in parallel by a state space model, configured to extract temporal features, and by a convolutional neural network, configured to extract spatial features. The resulting feature representations are combined to form a unified embedding, which is input to a prediction head to generate local and overall quality scores indicative of the perceptual quality of the video. In some embodiments, frame-level supervision is employed during training by comparing predicted per-frame scores to reference scores, improving accuracy and granularity. The invention enables robust, efficient, and scalable video quality assessment suitable for use with streaming optimization, compression, and quality monitoring systems, and is adaptable to various neural network backbones.

Background

FIELD OF THE INVENTION

The present invention relates generally to the field of automated video quality assessment. More particularly, the invention pertains to systems and methods for evaluating the perceptual quality of digital video content using machine learning techniques. BACKGROUND

The widespread growth of digital video content across streaming services, social media, video conferencing, and entertainment platforms has created an ongoing need for accurate and efficient assessment of video quality. As consumption of video media continues to increase, ensuring a high-quality viewing experience while optimizing bandwidth, storage, and processing resources has become an important objective for various content providers, service operators, and technology developers.

Traditionally, video quality assessment (VQA) has relied on objective, algorithmic metrics such as Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM). These full-reference metrics compare a compressed or processed video to an original, pristine reference version to quantify quality degradation. While these methods are computationally straightforward, they often exhibit poor correlation with subjective human perception, especially in complex or highly compressed video scenarios. Moreover, full-reference approaches are impractical for many real-world applications, such as user-generated content or live streaming, where reference videos are unavailable.

To ad

Claims

1. A computer-implemented method for assessing a quality of a digital video, comprising: receiving, by one or more processors, a digital video comprising a sequence of video frames; processing the video frames, by a hybrid neural network architecture comprising a first visual quality assessment (VQA) branch utilizing a state space model and a second VQA branch utilizing a convolutional neural network, wherein the first VQA branch extracts temporal features from the sequence of video frames by the state space model and the second VQA branch extracts spatial features from individual frames by the convolutional neural network; combining the temporal and spatial features to form a unified feature representation; and generating, by the neural network architecture using the unified feature representation, at least one quality score indicative of a perceptual quality of the digital video. || 15. A system comprising one or more processors and a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the system to: receive a digital video comprising a sequence of video frames; process the video frames, by a hybrid neural network architecture comprising a first visual quality assessment (VQA) branch utilizing a state space model and a second VQA branch utilizing a convolutional neural network, wherein the first VQA branch extracts temporal features from the sequence of video frames by the state space model and the second VQA branch extracts spatial features from individual frames by the convolutional neural network; combine the temporal and spatial features to form a unified feature representation; and generate by the neural network architecture using the unified feature representation, at least one quality score indicative of a perceptual quality of the digital video. || 16. The system set forth in claim 15 wherein the instructions stored in the computer-readable memory further cause the system to generate a quality score for each frame of the video, in addition to the at least one quality score for the sequence. || 17. The system set forth in claim 15 wherein the instructions stored in the computer-readable memory further cause the system to provide frame-level supervision during training by comparing predicted frame-level quality scores to reference scores generated by a pre-trained image quality assessment model, and transform the predicted frame-level quality scores via a learned mapping module to align with a distribution of the reference scores. || 18. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: receive a digital video comprising a sequence of video frames; process the video frames, by a hybrid neural network architecture comprising a first visual quality assessment (VQA) branch utilizing a state space model and a second VQA branch utilizing a convolutional neural network, wherein the first VQA branch extracts temporal features from the sequence of video frames by the state space model and the second VQA branch extracts spatial features from individual frames by the convolutional neural network; combine the temporal and spatial features to form a unified feature representation; and generate by the neural network architecture using the unified feature representation, at least one quality score indicative of a perceptual quality of the digital video. || 19. The non-transitory computer-readable medium set forth in claim 18 wherein the instructions, when executed by the one or more processors, cause the one or more processors to generate a quality score for each frame of the video, in addition to the at least one quality score for the sequence. || 20. The non-transitory computer-readable medium set forth in claim 18 wherein the instructions, when executed by the one or more processors, cause the one or more processors to provide frame-level supervision during training by comparing predicted frame-level quality scores to reference scores generated by a pre-trained image quality assessment model, and transform the predicted frame-level quality scores via a learned mapping module to align with a distribution of the reference scores.