- Number
- 20230045354
- Published
- 2023-02-09
- Filed
- 2021-08-06
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Farre Guiu; Miquel Angel et al.
- CPC
- G06F16/24573; G06F16/285; G06F16/75; G06N20/00
- Verdict
- Set aside content annotation taxonomy, business
- Source
- Google Patents · FreePatentsOnline
Abstract
According to one implementation, a system includes a computing platform having processing hardware, a system memory storing a software code; and a machine learning model based classifier. The processing hardware is configured to execute the software code to receive tagging quality assurance (QA) data including multiple terms applied as tags and corrections to those tags, to identify, using the tagging QA data, a first problematic term, and to classify, using the machine learning model based classifier, the first problematic term as one of confusing or flawed. The processing hardware is further configured to execute the software code to obtain, when the first problematic term is classified as confusing, a comparative sample for clarifying use of the first problematic term, and to obtain, when the first problematic term is classified as flawed, modification data for editing a predetermined annotation taxonomy including the first problematic term.
Background
BACKGROUND
Due to its popularity as a content medium, ever more video is being produced and made available to users. As a result, the efficiency with which video content can be annotated. i.e., “tagged,” and managed has become increasingly important to the producers, owners, and distributors of that video content. For example, annotation of video is an important part of the production process for television (TV) programming content and movies.
Tagging of video has traditionally been performed manually by human taggers, based on a predetermined set, or “taxonomy.” of terms that may be applied as tags, while quality assurance (QA) for the tagging process is typically performed by human QA reviewers. However, in a typical video production environment, there may be such a large number of videos to be annotated that manual tagging and review become impracticable. In response, various automated systems for performing content tagging and QA review have been developed or are in development. While offering efficiency advantages over traditional manual techniques, the performance of automated systems, like the performance of human taggers, depends to a significant extent on the relevance and specificity of the typically closed set of terms included in the annotation taxonomy. Consequently, there is a need in the art for systems and methods for enhancing the performance of automated and human taggers alike through the performance-based evolution of content annotation taxon
Claims
1. A system comprising: a computing platform including a processing hardware and a system memory storing a software code; and a machine learning model based classifier, the processing hardware configured to execute the software code to: receive tagging quality assurance (QA) data including a plurality of terms applied as tags and a plurality of corrections to the tags; identify, using the tagging QA data, a first problematic term of the plurality of terms; classify, using the machine learning model based classifier, the first problematic term as one of a confusing term or a flawed term; obtain, when the first problematic term is classified as the confusing term, a comparative sample for clarifying use of the first problematic term as a tag; and obtain, when the first problematic term is classified as the flawed term, a modification data for editing a predetermined annotation taxonomy including the first problematic term. ||
10. A method for use by a system including a computing platform having a processing hardware and a system memory storing a software code and a machine learning model based classifier, the method comprising: receiving, by the software code executed by the processing hardware, tagging quality assurance (QA) data including a plurality of terms applied as tags and a plurality of corrections to the tags; identifying, by the software code executed by the processing hardware and using the tagging QA data, a first problematic term of the plurality of terms; classifying, by the software code executed by the processing hardware and using the machine learning model based classifier, the first problematic term as one of a confusing term or a flawed term; obtaining, by the software code executed by the processing hardware when the first problematic term is classified as the confusing term, a comparative sample for clarifying use of the first problematic term as a tag; and obtaining, by the software code executed by the processing hardware when the first problematic term is classified as the flawed term, a modification data for editing a predetermined annotation taxonomy including the first problematic term. ||
16. A method for use by a system including a computing platform having a processing hardware and a system memory storing a software code, the method comprising: receiving, by the software code executed by the processing hardware, tagging quality assurance (QA) data including a plurality of terms applied as tags by a trained machine learning model based automated tagging system, and a plurality of corrections to the tags; identifying, by the software code executed by the processing hardware and using the tagging QA data, a first automated problematic term of the plurality of terms; classifying, by the software code executed by the processing hardware and using the tagging QA data, the first automated problematic term as one of a re-trainable term or a flawed term; obtaining, by the software code executed by the processing hardware when the first automated problematic term is classified as the re-trainable term, one or more parameters for adjusting the trained machine learning model based automated tagging system; and obtaining, by the software code executed by the processing hardware when first automated problematic term is classified as the flawed term, a modification data for editing a predetermined annotation taxonomy including the flawed term.