Outer Rim Archives
Archives · 2022 · 20220292116

Application (pre-grant publication)

Constrained Multi-Label Dataset Partitioning for Automated Machine Learning

Number
20220292116
Published
2022-09-15
Filed
2021-03-09
Assignee
Disney Enterprises, Inc.
Inventors
Martin; Marc Junyent, Farre Guiu; Miquel Angel
CPC
G06F16/285; G06F18/214; G06N20/00; G06F16/278; G06F18/217
Verdict
Set aside generic AutoML data engineering
Source
Google Patents · FreePatentsOnline

Abstract

A system includes a computing platform having processing hardware and a memory storing a software code. The processing hardware executes the software code to receive a dataset including at least some data samples having multiple metadata labels, and identify a partitioning constraint and a partitioning of the dataset into data subsets. The software code also executed obtains, for each metadata label, a desired distribution ratio based on the number of the data subsets and a total number of instances that each metadata label has been applied to the data samples, aggregates, using the partitioning constraint, the data samples into data sample groups, assigns, using the partitioning constraint and the desired distribution ratio for each of the metadata labels, each of the data sample groups to one of the data subsets, wherein each of the data subsets are unique, and trains, using one of the data subsets, a machine learning model.

Background

BACKGROUND

Media content in the form of video, audio, and text, for example, is continuously being produced and made available to users whose appetite for such content is nearly limitless. As a result, the efficiency with which media content can be annotated and managed, as well as the accuracy with which annotations are applied to the media content, have become increasingly important to the producers and owners of that content.

For example, annotation of video is an important part of the production process for television (TV) programming and movies, and is typically performed manually by human annotators. However, such manual annotation, or “tagging,” of video is a labor intensive and time consuming process. Moreover, in a typical video production environment there may be such a large number of videos to be annotated that manual tagging becomes impracticable. In response, automated solutions for annotating media content have been developed. While offering efficiency advantages over traditional manual tagging, automated annotation systems are typically more prone to error than human taggers. In order to improve the accuracy of automated annotation systems, it is desirable to ensure that the data used to train such systems does not overlap with data used to validate those systems for use. Consequently, there is a need in the art for an automated solution for appropriately partitioning datasets for use in training and validation of automated annotation systems.

Claims

1. A system comprising: a computing platform including a processing hardware and a system memory storing a software code; the processing hardware configured to execute the software code to: receive a dataset including a plurality of data samples, at least some of the plurality of data samples having a plurality of metadata labels applied thereto; identify a partitioning constraint and a partitioning of the dataset into a plurality of data subsets; obtain, for each of the plurality of metadata labels, a desired distribution ratio based on the plurality of data subsets and a total number of instances that each of the plurality of metadata labels has been applied to the plurality of data samples; aggregate, using the partitioning constraint, the plurality of data samples into a plurality of data sample groups; assign, using the partitioning constraint and the desired distribution ratio for each respective one of the plurality of metadata labels, each of the plurality of data sample groups to one of the plurality of data subsets, wherein each of the plurality of data subsets are unique; and train, using one of the plurality of data subsets, a machine learning model. || 0. || 11. A method for use by a system including a computing platform having a processing hardware and a system memory storing a software code, the method comprising: receiving, by the software code executed by the processing hardware, a dataset including a plurality of data samples, at least some of the plurality of data samples having a plurality of metadata labels applied thereto; identifying, by the software code executed by the processing hardware, a partitioning constraint and a partitioning of the dataset into a plurality of data subsets; obtaining, by the software code executed by the processing hardware for each of the plurality of metadata labels, a desired distribution ratio based on the plurality of data subsets and a total number of instances that each of the plurality of metadata labels has been applied to the plurality of data samples; aggregating, by the software code executed by the processing hardware and using the partitioning constraint, the plurality of data samples into a plurality of data sample groups; assigning, by the software code executed by the processing hardware and using the partitioning constraint and the desired distribution ratio for each respective one of the plurality of metadata labels, each of the plurality of data sample groups to one of the plurality of data subsets, wherein each of the data subsets are unique; and training, by the software code executed by the processing hardware and using one of the plurality of data subsets, a machine learning model. || 0.