Outer Rim Archives
Archives · 2025 · 12222965

Granted patent

Constrained multi-label dataset partitioning for automated machine learning

Number
12222965
Published
2025-02-11
Filed
2021-03-09
Assignee
Disney Enterprises, Inc.
Inventors
Martin; Marc Junyent et al.
CPC
G06F18/217; G06F16/285; G06N20/00; G06F16/278; G06F18/214
Verdict
Set aside generic ML data-partitioning research
Source
Google Patents · FreePatentsOnline

Abstract

A system includes a computing platform having processing hardware and a memory storing a software code. The processing hardware executes the software code to receive a dataset including at least some data samples having multiple metadata labels, and identify a partitioning constraint and a partitioning of the dataset into data subsets. The software code also executed obtains, for each metadata label, a desired distribution ratio based on the number of the data subsets and a total number of instances that each metadata label has been applied to the data samples, aggregates, using the partitioning constraint, the data samples into data sample groups, assigns, using the partitioning constraint and the desired distribution ratio for each of the metadata labels, each of the data sample groups to one of the data subsets, wherein each of the data subsets are unique, and trains, using one of the data subsets, a machine learning model.

Background

BACKGROUND (1) Media content in the form of video, audio, and text, for example, is continuously being produced and made available to users whose appetite for such content is nearly limitless. As a result, the efficiency with which media content can be annotated and managed, as well as the accuracy with which annotations are applied to the media content, have become increasingly important to the producers and owners of that content. (2) For example, annotation of video is an important part of the production process for television (TV) programming and movies, and is typically performed manually by human annotators. However, such manual annotation, or “tagging,” of video is a labor intensive and time consuming process. Moreover, in a typical video production environment there may be such a large number of videos to be annotated that manual tagging becomes impracticable. In response, automated solutions for annotating media content have been developed. While offering efficiency advantages over traditional manual tagging, automated annotation systems are typically more prone to error than human taggers. In order to improve the accuracy of automated annotation systems, it is desirable to ensure that the data used to train such systems does not overlap with data used to validate those systems for use. Consequently, there is a need in the art for an automated solution for appropriately partitioning datasets for use in training and validation of automated annotation systems.

Claims

1. A system comprising: a computing platform including a processing hardware and a system memory storing a software code; the processing hardware configured to execute the software code to: receive, as part of an automated process, a dataset including a plurality of data samples, wherein each of at least some of the plurality of data samples has a plurality of metadata labels applied thereto, the dataset further including metadata specifying (i) how many data subsets the dataset is to be partitioned into, and (ii) a partitioning constraint for partitioning the dataset into a plurality of data subsets; identify, as part of the automated process, the partitioning constraint and a partitioning of the dataset into the plurality of data subsets; obtain, as part of the automated process, for each of the plurality of metadata labels, a desired distribution ratio based on the plurality of data subsets and a total number of instances that each of the plurality of metadata labels has been applied to the plurality of data samples; aggregate, as part of the automated process and subject to the partitioning constraint, the plurality of data samples into a plurality of data sample groups; assign, as part of the automated process, subject to the partitioning constraint and using the desired distribution ratio for each respective one of the plurality of metadata labels, each of the plurality of data sample groups to one of the plurality of data subsets, wherein each of the plurality of data subsets are unique; train, as part of the automated process and using one of the plurality of data subsets, a machine learning model to provide a trained machine learning model; and annotate media content, using the trained machine learning model. || 0. || 11. A method for use by a system including a computing platform having a processing hardware and a system memory storing a software code, the method comprising: receiving, by the software code executed by the processing hardware as part of an automated process, a dataset including a plurality of data samples, wherein each of at least some of the plurality of data samples has a plurality of metadata labels applied thereto, the dataset further including metadata specifying (i) how many data subsets the dataset is to be partitioned into, and (ii) a partitioning constraint for partitioning the dataset into a plurality of data subsets; identifying, by the software code executed by the processing hardware as part of the automated process, the partitioning constraint and a partitioning of the dataset into the plurality of data subsets; obtaining, by the software code executed by the processing hardware for each of the plurality of metadata labels as part of the automated process, a desired distribution ratio based on the plurality of data subsets and a total number of instances that each of the plurality of metadata labels has been applied to the plurality of data samples; aggregating, by the software code executed by the processing hardware as part of the automated process, subject to the partitioning constraint, the plurality of data samples into a plurality of data sample groups; assigning, by the software code executed by the processing hardware as part of the automated process, subject to the partitioning constraint and using the desired distribution ratio for each respective one of the plurality of metadata labels, each of the plurality of data sample groups to one of the plurality of data subsets, wherein each of the data subsets are unique; training, by the software code executed by the processing hardware as part of the automated process and using one of the plurality of data subsets, a machine learning model to provide a trained machine learning model; and annotating media content, using the trained machine learning model. || 0.