Application (pre-grant publication)
SEMIAUTOMATIC MACHINE LEARNING MODEL IMPROVEMENT AND BENCHMARKING
- Number
- 20190034822
- Published
- 2019-01-31
- Filed
- 2017-07-27
- Assignee
- Disney Enterprises, Inc.
- Inventors
- FARRÉ GUIU; Miquel Angel et al.
- CPC
- G06N5/022; G06F18/2178; G06V10/776; G06V10/774; G06F18/217; G06V20/41; G06F18/2155; G06F18/214; G06V10/7784; G06N20/00
- Verdict
- Set aside semiautomatic ML model improvement/benchmarking, generic ML tooling
- Source
- Google Patents · FreePatentsOnline
Abstract
Systems, methods, and articles of manufacture to perform an operation comprising processing, by a machine learning (ML) algorithm and a ML model, a plurality of images in a first dataset, wherein the ML model was generated based on a plurality of images in a training dataset, receiving user input reviewing a respective set of tags applied to each image in the first data set as a result of the processing, identifying, based on a first confusion matrix generated based on the user input and the sets of tags applied to the images in the first data set, a first labeling error in the training dataset, determining a type of the first labeling error based on a second confusion matrix, and modifying the training dataset based on the determined type of the first labeling error.
Background
BACKGROUNDField of the Invention
Embodiments disclosed herein relate to machine learning. More specifically, embodiments disclosed herein relate to semiautomatic machine learning model improvement and benchmarking.Description of the Related Art
Many vendors offer machine learning (ML) algorithms to identify different types of elements in media files at high levels of accuracy. However, to provide high levels of accuracy, the algorithms must be trained based on a training dataset. Nevertheless, preparing an accurate training dataset to train the ML algorithms is difficult due to the need of keeping the dataset updated (e.g., cleaning the dataset, correcting errors in the dataset, adding more data, and the like). Furthermore, updating the dataset does not provide a description of the improvement or saturation of the ML algorithms after adding new edge cases.SUMMARY
In one embodiment, a method comprises processing, by a machine learning (ML) algorithm and a ML model, a plurality of images in a first dataset, wherein the ML model was generated based on a plurality of images in a training dataset, receiving user input reviewing a respective set of tags applied to each image in the first data set as a result of the processing, identifying, based on a first confusion matrix generated based on the user input and the sets of tags applied to the images in the first data set, a first labeling error in the training dataset, determining a type of the first labeling error