Outer Rim Archives
Archives · 2023 · 20230068502

Application (pre-grant publication)

Multi-Modal Content Based Automated Feature Recognition

Number
20230068502
Published
2023-03-02
Filed
2021-08-30
Assignee
Disney Enterprises, Inc.
Inventors
Pernias; Pablo et al.
CPC
G06N20/00; G06F16/686; G06F16/75; G06F16/735; G06N3/0455; G06V10/454; G06V20/46; G06F16/7867; G06F18/253; G06N3/045; G06F16/45; G06N3/09; G06F18/213
Verdict
Set aside generic content feature recognition, business
Source
Google Patents · FreePatentsOnline

Abstract

A system includes a computing platform having processing hardware, and a memory storing software code and a machine learning (ML) model-based feature classifier. When executed, the software code receives media content including a first media component corresponding to a first media mode and a second media component corresponding to a second media mode, encodes the first media component using a first encoder to generate multiple first embedding vectors, and encodes the second media component using a second encoder to generate multiple second embedding vectors. The software code further combines the first embedding vectors and the second embedding vectors to provide an input data structure for a neural network mixer, process, using the neural network mixer, the input data structure to provide feature data corresponding to a feature of the media content, and predict, using the ML model-based feature classifier and the feature data, a classification of the feature.

Background

BACKGROUND

Due to its popularity as a content medium, ever more audio-video (AV) content is being produced and made available to users. As a result, the efficiency with which such AV content can be annotated. i.e., “tagged.” and managed has become increasingly important to the producers, owners, and distributors of that content. For example, annotation of AV content is an important part of the production process for television (TV) programming content and movies.

Tagging of AV content has traditionally been performed manually by human taggers. However, in a typical AV content production environment, there may be so much content to be annotated that manual tagging becomes impracticable. In response, various automated systems for performing content tagging have been developed or are in development. While offering efficiency advantages over traditional manual techniques for the identification of features that can be recognized based on a single content mode, such as video alone, audio alone, or text alone, conventional automated tagging solutions can be unreliable when used to classify features requiring more than one content mode for their identification. Consequently, there is a need in the art for systems and methods enabling reliable multi-modal content based automated feature recognition.

Claims

1. A system comprising: a computing platform including a processing hardware and a system memory storing a software code and a machine learning (ML) model-based feature classifier; the processing hardware configured to execute the software code to: receive media content including a first media component corresponding to a first media mode and a second media component corresponding to a second media mode; encode the first media component, using a first encoder of the software code, to generate a first plurality of embedding vectors; encode the second media component, using a second encoder of the software code, to generate a second plurality of embedding vectors; combine the first plurality of embedding vectors and the second plurality of embedding vectors to provide an input data structure for a neural network mixer of the software code; process, using the neural network mixer, the input data structure to provide a feature data corresponding to a feature of the media content; and predict, using the ML model-based feature classifier and the feature data, a classification of the feature. || 10. A method for use by a system including a computing platform having a processing hardware, and a system memory storing a software code and a machine learning (ML) model-based feature classifier, the method comprising: receiving, by the software code executed by the processing hardware, media content including a first media component corresponding to a first media mode and a second media component corresponding to a second media mode; encoding, by the software code executed by the processing hardware and using a first encoder of the software code, the first media component, to generate a first plurality of embedding vectors; encoding, by the software code executed by the processing hardware and using a second encoder of the software code, the second media component, to generate a second plurality of embedding vectors; combining, by the software code executed by the processing hardware, the first plurality of embedding vectors and the second plurality of embedding vectors to provide an input data structure for a neural network mixer of the software code; processing, by the software code executed by the processing hardware and using the neural network mixer, the input data structure to provide a feature data corresponding to a feature of the media content; and predicting, by the software code executed by the processing hardware and using the ML model-based feature classifier and the feature data, a classification the feature. || 19. A method for use by a system including a computing platform having a processing hardware, and a system memory storing a feature locator software code, the method comprising: receiving, by the feature locator software code executed by the processing hardware, a feature data corresponding to a feature of media content; receiving, by the feature locator software code executed by the processing hardware, a classification of the feature; and identifying, by the feature locator software code executed by the processing hardware based on the classification and the feature data, a timestamp of the media content corresponding to the feature.