Outer Rim Archives
Archives · 2021 · 10891985

Granted patent

Bi-level specificity content annotation using an artificial neural network

Number
10891985
Published
2021-01-12
Filed
2019-07-11
Assignee
Disney Enterprises, Inc.
Inventors
Farre Guiu; Miquel Angel, Alfaro Vendrell; Monica, Aparicio Isarn; Albert, Fojo; Daniel, Martin; Marc Junyent, Accardo; Anthony M., Swerdlow; Avner
CPC
G06N3/045; G06N3/0455; G06N3/0464; G06N3/08; G06N3/09; G06V20/46; G11B27/031; G11B27/19; G11B27/34
Verdict
Set aside content annotation, business
Source
Google Patents · FreePatentsOnline

Abstract

A content annotation system includes a computing platform having a hardware processor and a memory storing a tagging software code including an artificial neural network (ANN). The hardware processor executes the tagging software code to receive content having a content interval including an image of a generic content feature, encode the image into a latent vector representation of the image using an encoder of the ANN, and use a first decoder of the ANN to generate a first tag describing the generic content feature based on the latent vector representation. When a specific content feature learned by the ANN corresponds to the generic content to feature described by the first tag, the tagging software code uses a second decoder of the ANN to generate a second tag uniquely identifying the specific content feature based on the latent vector representation, and tags the content interval with the first and second tags.

Background

BACKGROUND (1) Due to its nearly universal popularity as a content medium, ever more video is being produced and made available to users. As a result, the efficiency with which video content can be annotated and managed has become increasingly important to the producers and owners of that video content. (2) Annotation of video content has traditionally been performed manually by human annotators. However, such manual annotation, or 'tagging,' of video is a labor intensive and time consuming process. Consequently, there is a need in the art for an automated solution for annotating content that substantially minimizes the amount of content, such as video, that needs to be manually processed SUMMARY (3) There are provided systems and methods for automating the performance of bi-level specificity content annotation using an artificial neural network (ANN), substantially as shown in and/or described in connection with at least one of the figures, and as set forth more completely in the claims.

Claims

1. A content annotation system comprising: a computing platform having a hardware processor and a memory storing a tagging software code including an artificial neural network (ANN) trained to identify at least one generic content feature and at least one specific content feature; the hardware processor configured to execute the tagging software code to: receive content having a content interval including an image of the at least one generic content feature; encode the image into a latent vector representation of the image using an encoder of the ANN; generate a first tag describing the at least one generic content feature based on the latent vector representation, using a first decoder of the ANN; when the at least one specific content feature corresponds to the at least one generic content feature described by the first tag, generate a second tag uniquely identifying the at least one specific content feature based on the latent vector representation, using a second decoder of the ANN; and tag the content interval with the first tag and the second tag. || 11. A method for use by a content annotation system including a computing platform having a hardware processor and a memory storing a software code including an artificial neural network (ANN) trained to identify at least one generic content feature and at least one specific content feature, the method comprising: receiving, by the software code executed by the hardware processor, content having a content interval including an image of the at least one generic content feature; encoding the image into a latent vector representation of the image, by the software code executed by the hardware processor, and using an encoder of the ANN; generating, by the software code executed by the hardware processor, a first tag describing the at least one generic content feature based on the latent vector representation, using a first decoder of the ANN; when the at least one specific content feature corresponds to the at least one generic content feature described by the first tag, generating, by the software code executed by the hardware processor, a second tag uniquely identifying the at least one specific content features based on the latent vector representation, using a second decoder of the ANN; and tagging the content interval with the first tag and the second tag, by the software code executed by the hardware processor.