Outer Rim Archives
Archives · 2020 · 20200151459

Application (pre-grant publication)

Guided Training for Automation of Content Annotation

Number
20200151459
Published
2020-05-14
Filed
2019-03-13
Assignee
Disney Enterprises, Inc.
Inventors
Farre Guiu; Miquel Angel, Petrillo; Matthew, Alfaro Vendrell; Monica, Junyent Martin; Marc, Fojo; Daniel, Accardo; Anthony M., Swerdlow; Avner, Navarre; Katharine
CPC
G06F18/214; G06F18/28; G06F18/41; G06T7/60; G06T7/70; G06V10/7784; G06V10/945; G06V20/41; G06V20/47
Verdict
Set aside content annotation ML tooling, business
Source
Google Patents · FreePatentsOnline

Abstract

According to one implementation, a system for automating content annotation includes a computing platform having a hardware processor and a system memory storing an automation training software code. The hardware processor executes the automation training software code to initially train a content annotation engine using labeled content, test the content annotation engine using a first test set of content obtained from a training database, and receive corrections to a first automatically annotated content set resulting from the test. The hardware processor further executes the automation training software code to further train the content annotation engine based on the corrections, determine one or more prioritization criteria for selecting a second test set of content for testing the content annotation engine based on the statistics relating to the first automatically annotated content, and select the second test set of content from the training database based on the prioritization criteria.

Background

BACKGROUND

Due to its nearly universal popularity as a content medium, ever more video is being produced and made available to users. As a result, the efficiency with which video content can be annotated and managed has become increasingly important to the producers of that video content.

For example, annotation of video is an important part of the production process for television (TV) programming and movies, and is typically performed manually by human annotators. However, such manual annotation, or “tagging”, of video is a labor intensive and time consuming process. Moreover, in a typical video production environment there may be such a large number of videos to be annotated that manual tagging becomes impracticable. Consequently, there is a need in the art for an automated solution for annotating content that substantially minimizes the amount of content, such as video, that needs to be manually processed.SUMMARY

There are provided systems and methods for automating content annotation, substantially as shown in and/or described in connection with at least one of the figures, and as set forth more completely in the claims.

Claims

1. A system for automating content annotation, the system comprising: a computing platform including a hardware processor and a system memory; an automation training software code stored in the system memory; the hardware processor configured to execute the automation training software code to: initially train a content annotation engine using a labeled content; test the content annotation engine using a first test set of content obtained from a training database, resulting in a first automatically annotated content set; receive a plurality of corrections to the first automatically annotated content set; further train the content annotation engine based on the plurality of corrections to the first automatically annotated content set; determine at least one prioritization criteria for selecting a second test set of content for testing the content annotation engine based on statistics relating to the first automatically annotated content set; and select the second test set of content from the training database based on the at least one prioritization criteria. 12. A method for use by a system for automating content annotation that includes a computing platform having a hardware processor and a system memory storing an automation training software code, the method comprising: initially training, by the automation training software code executed by the hardware processor, a content annotation engine using a labeled content; testing, by the automation training software code executed by the hardware processor, the content annotation engine using a first test set of content obtained from a training database, resulting in a first automatically annotated content set; receiving, by the automation training software code executed by the hardware processor, a plurality of corrections to the first automatically annotated content set; further training, by the automation training software code executed by the hardware processor, the content annotation engine based on the plurality of corrections to the first automatically annotated content set; determining, by the automation training software code executed by the hardware processor, at least one prioritization criteria for selecting a second test set of content for testing the content annotation engine based on statistics relating to the first automatically annotated content set; and selecting, by the automation training software code executed by the hardware processor, the second test set of content from the training database based on the at least one prioritization criteria. 23. A computer-readable non-transitory medium having stored thereon an automation training software code, which when executed by a hardware processor, instantiates a method comprising: initially training a content annotation engine using a labeled content; testing the content annotation engine using a first test set of content obtained from a training database, resulting in a first automatically annotated content set; receiving a plurality of corrections to the first automatically annotated content set; further training the content annotation engine based on the plurality of corrections to the first automatically annotated content set; determining at least one prioritization criteria for selecting a second test set of content for testing the content annotation engine based on statistics relating to the first automatically annotated content set; and selecting the second test set of content from the training database based on the at least one prioritization criteria.