Application (pre-grant publication)
Machine Learning Model-Based Content Anonymization
- Number
- 20220343020
- Published
- 2022-10-27
- Filed
- 2022-03-24
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Farre Guiu; Miquel Angel, Pernias; Pablo, Martin; Marc Junyent
- CPC
- G06N3/045; G06N3/08; G06F21/6254; G06N3/096
- Verdict
- Set aside content anonymization, privacy/business
- Source
- Google Patents · FreePatentsOnline
Abstract
A system includes a computing platform having processing hardware, and a system memory storing software code and a machine learning (ML) model. The processing hardware is configured to execute the software code to receive from a client, a request for a dataset, the request identifying a content type of the dataset, obtain the dataset having the content type, and select, based on the content type, an anonymization technique for the dataset, the anonymization technique selected so as to render at least one feature included in the dataset recognizable but unidentifiable. The processing hardware is further configured to execute the software code to anonymize, using the ML model and the selected anonymization technique, the at least one feature included in the dataset, and to output to the client, in response to the request, an anonymized dataset including the at least one anonymized feature.
Background
BACKGROUND
There are many undertakings that can be advanced more efficiently when performed collaboratively. However, due to any of potentially myriad privacy or confidentiality concerns, it may be undesirable to share data including proprietary or otherwise sensitive information in a collaborative or other semi-public or public environment.
By way of example, collaborative machine learning development endeavors, such as “hackathons” for instance, can advantageously accelerate the process of identifying and optimizing machine learning models for use in a variety of applications, such as activity recognition, location recognition, facial recognition, and object recognition. However, due to the proprietary nature or sensitivity of certain types of content, it may be undesirable to make such content generally available for use in model training.
One conventional approach to satisfying the competing interests of content availability and content security for machine learning development is to utilize a remote execution platform as a privacy shield between model developers and content owners. According to this approach, the remote execution platform can mediate training of a machine learning model using proprietary content, while sequestering that content from the model developer. However, one disadvantage of this approach is that, because the model developer is prevented from accessing the content used for training, it is difficult or impossible to accurately