Outer Rim Archives
Archives · 2019 · 20190034822

Application (pre-grant publication)

SEMIAUTOMATIC MACHINE LEARNING MODEL IMPROVEMENT AND BENCHMARKING

Number
20190034822
Published
2019-01-31
Filed
2017-07-27
Assignee
Disney Enterprises, Inc.
Inventors
FARRÉ GUIU; Miquel Angel et al.
CPC
G06N5/022; G06F18/2178; G06V10/776; G06V10/774; G06F18/217; G06V20/41; G06F18/2155; G06F18/214; G06V10/7784; G06N20/00
Verdict
Set aside semiautomatic ML model improvement/benchmarking, generic ML tooling
Source
Google Patents · FreePatentsOnline

Abstract

Systems, methods, and articles of manufacture to perform an operation comprising processing, by a machine learning (ML) algorithm and a ML model, a plurality of images in a first dataset, wherein the ML model was generated based on a plurality of images in a training dataset, receiving user input reviewing a respective set of tags applied to each image in the first data set as a result of the processing, identifying, based on a first confusion matrix generated based on the user input and the sets of tags applied to the images in the first data set, a first labeling error in the training dataset, determining a type of the first labeling error based on a second confusion matrix, and modifying the training dataset based on the determined type of the first labeling error.

Background

BACKGROUNDField of the Invention

Embodiments disclosed herein relate to machine learning. More specifically, embodiments disclosed herein relate to semiautomatic machine learning model improvement and benchmarking.Description of the Related Art

Many vendors offer machine learning (ML) algorithms to identify different types of elements in media files at high levels of accuracy. However, to provide high levels of accuracy, the algorithms must be trained based on a training dataset. Nevertheless, preparing an accurate training dataset to train the ML algorithms is difficult due to the need of keeping the dataset updated (e.g., cleaning the dataset, correcting errors in the dataset, adding more data, and the like). Furthermore, updating the dataset does not provide a description of the improvement or saturation of the ML algorithms after adding new edge cases.SUMMARY

In one embodiment, a method comprises processing, by a machine learning (ML) algorithm and a ML model, a plurality of images in a first dataset, wherein the ML model was generated based on a plurality of images in a training dataset, receiving user input reviewing a respective set of tags applied to each image in the first data set as a result of the processing, identifying, based on a first confusion matrix generated based on the user input and the sets of tags applied to the images in the first data set, a first labeling error in the training dataset, determining a type of the first labeling error

Claims

1. A method, comprising: processing, by a machine learning (ML) algorithm and a ML model, a plurality of images in a first dataset, wherein the ML model was generated based on a plurality of images in a training dataset; receiving user input reviewing a respective set of tags applied to each image in the first data set as a result of the processing; identifying, based on a first confusion matrix generated based on the user input and the sets of tags applied to the images in the first data set, a first labeling error in the training dataset; determining a type of the first labeling error based on a second confusion matrix; and modifying the training dataset based on the determined type of the first labeling error. 8. The computer program product, comprising: a computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code executable by a processor to perform an operation comprising: processing, by a machine learning (ML) algorithm and a ML model, a plurality of images in a first dataset, wherein the ML model was generated based on a plurality of images in a training dataset; receiving user input reviewing a respective set of tags applied to each image in the first data set as a result of the processing; identifying, based on a first confusion matrix generated based on the user input and the sets of tags applied to the images in the first data set, a first labeling error in the training dataset; determining a type of the first labeling error based on a second confusion matrix; and modifying the training dataset based on the determined type of the first labeling error. 15. A system, comprising: one or more computer processors; and a memory containing a program which when executed by the processors performs an operation comprising: processing, by a machine learning (ML) algorithm and a ML model, a plurality of images in a first dataset, wherein the ML model was generated based on a plurality of images in a training dataset; receiving user input reviewing a respective set of tags applied to each image in the first data set as a result of the processing; identifying, based on a first confusion matrix generated based on the user input and the sets of tags applied to the images in the first data set, a first labeling error in the training dataset; determining a type of the first labeling error based on a second confusion matrix; and modifying the training dataset based on the determined type of the first labeling error.