Outer Rim Archives
Archives · 2025 · 12461967

Granted patent

Multi-modal content based automated feature recognition

Number
12461967
Published
2025-11-04
Filed
2021-08-30
Assignee
Disney Enterprises Inc.
Inventors
Pernias; Pablo et al.
CPC
G06F18/253; G06V10/454; G06F16/686; G06F18/213; G06F16/735; G06F16/7867; G06F16/45; G06N3/045; G06N20/00; G06V20/46; G06F16/75; G06N3/0455; G06N3/09
Verdict
Set aside content feature recognition/indexing, business
Source
Google Patents · FreePatentsOnline

Abstract

A system includes a computing platform having processing hardware, and a memory storing software code and a machine learning (ML) model-based feature classifier. When executed, the software code receives media content including a first media component corresponding to a first media mode and a second media component corresponding to a second media mode, encodes the first media component using a first encoder to generate multiple first embedding vectors, and encodes the second media component using a second encoder to generate multiple second embedding vectors. The software code further combines the first embedding vectors and the second embedding vectors to provide an input data structure for a neural network mixer, process, using the neural network mixer, the input data structure to provide feature data corresponding to a feature of the media content, and predict, using the ML model-based feature classifier and the feature data, a classification of the feature.

Background

BACKGROUND (1) Due to its popularity as a content medium, ever more audio-video (AV) content is being produced and made available to users. As a result, the efficiency with which such AV content can be annotated. i.e., “tagged.” and managed has become increasingly important to the producers, owners, and distributors of that content. For example, annotation of AV content is an important part of the production process for television (TV) programming content and movies. (2) Tagging of AV content has traditionally been performed manually by human taggers. However, in a typical AV content production environment, there may be so much content to be annotated that manual tagging becomes impracticable. In response, various automated systems for performing content tagging have been developed or are in development. While offering efficiency advantages over traditional manual techniques for the identification of features that can be recognized based on a single content mode, such as video alone, audio alone, or text alone, conventional automated tagging solutions can be unreliable when used to classify features requiring more than one content mode for their identification. Consequently, there is a need in the art for systems and methods enabling reliable multi-modal content based automated feature recognition.

Claims

1. A system comprising: a computing platform including a processing hardware and a system memory storing a software code and a machine learning (ML) model-based feature classifier; the processing hardware configured to execute the software code to: receive media content including a video component corresponding to a video mode and one of a text component or an audio component corresponding to one of a text mode or an audio mode; encode a plurality of video frames of the video component, using a first encoder of the software code, to generate a plurality of video embedding vectors; encode the one of the text component or the audio component, using a second encoder of the software code, to generate a plurality of audio embedding vectors or a plurality of text embedding vectors; combine the plurality of video embedding vectors and one of the plurality of audio embedding vectors or the plurality of text embedding vectors to provide an input data structure for a neural network mixer of the software code; process, using the neural network mixer, the input data structure to provide a feature data corresponding to a feature of the media content, wherein the neural network mixer is tuned to provide the feature data such that the feature data preferentially focuses on one of objects, locations, performers, characters, or activities depicted in the video component over others of the objects, locations, performers, characters, or activities depicted in the video component; and predict, using the ML model-based feature classifier and the feature data, a classification of the feature. || 11. A method for use by a system including a computing platform having a processing hardware, and a system memory storing a software code and a machine learning (ML) model-based feature classifier, the method comprising: receiving, by the software code executed by the processing hardware, media content including a video component corresponding to a video mode and one of a text component or an audio component corresponding to one of a text mode or an audio mode; encoding, by the software code executed by the processing hardware and using a first encoder of the software code, a plurality of video frames of the video component, to generate a plurality of video embedding vectors; encoding, by the software code executed by the processing hardware and using a second encoder of the software code, the one of the text component or the audio component, to generate a plurality of audio embedding vectors or a plurality of text embedding vectors; combining, by the software code executed by the processing hardware, the plurality of video embedding vectors and one of the plurality of audio embedding vectors or the plurality of text embedding vectors to provide an input data structure for a neural network mixer of the software code; processing, by the software code executed by the processing hardware and using the neural network mixer, the input data structure to provide a feature data corresponding to a feature of the media content, wherein the neural network mixer is tuned to provide the feature data such that the feature data preferentially focuses on one of objects, locations, performers, characters, or activities depicted in the video component over others of the objects, locations, performers, characters, or activities depicted in the video component; and predicting, by the software code executed by the processing hardware and using the ML model-based feature classifier and the feature data, a classification the feature. || 19. A method for use by a system including a computing platform having a processing hardware, and a system memory storing a feature locator software code, the method comprising: receiving, by the feature locator software code executed by the processing hardware, a feature data corresponding to a feature of media content, wherein the feature data is provided by a neural network mixer based on an input data structure including a plurality of video embedding vectors in combination with one of a plurality of audio embedding vectors or a plurality of text embedding vectors, wherein the neural network mixer is tuned to provide the feature data such that the feature data preferentially focuses on one of objects, locations, performers, characters, or activities depicted in a video component of the media content over others of the objects, locations, performers, characters, or activities depicted in the video component; receiving, by the feature locator software code executed by the processing hardware, a classification of the feature; and identifying, by the feature locator software code executed by the processing hardware, based on the classification and the feature data and using a feature mask comprising a matrix having a same dimensionality as the feature data, a timestamp of the media content corresponding to the feature.