Outer Rim Archives
Archives · 2022 · 11314706

Granted patent

Metadata aggregation using a trained entity matching predictive model

Number
11314706
Published
2022-04-26
Filed
2020-08-21
Assignee
Disney Enterprises, Inc.
Inventors
Stoafer; Christopher C., Pujol; Jordi Badia, Bravo; Francesc Josep Guitart, Martin; Marc Junyent, Farre Guiu; Miquel Angel, Lawson; Calvin, Luerken; Erick L.
CPC
G06F11/3419; G06F11/3433; G06F16/1794; G06F16/215; G06F16/217; G06F16/24542; G06F16/24549; G06F16/2462; G06F16/3346; G06F16/38; G06F16/383; G06F17/11; G06F17/18; G06F30/27; G06F9/4881; G06N7/01
Verdict
Set aside metadata aggregation, data-engineering business
Source
Google Patents · FreePatentsOnline

Abstract

A metadata aggregation system includes a computing platform having a hardware processor and a memory storing a software code including a trained entity matching predictive model trained using training data obtained from a reference database. The hardware processor executes the software code to obtain metadata inputs from multiple sources, conform the metadata inputs to a common format, match, using the trained entity matching predictive model, at least some of the conformed metadata inputs to the same entity, and determine, using the trained entity matching predictive model, a confidence score for each match. The software code further sends a request to one or more human editor(s) for confirmation of each match having a confidence score greater than a first threshold and less than a second threshold, and updates the reference database, in response to receiving a confirmation that at least one match is a confirmed match, to include the confirmed match.

Background

BACKGROUND (1) Popular movies, television programs, sports teams, and other pop culture “entities” are typically the subjects of extensive commentary from a variety of different sources. For example, a movie or movie franchise may be the subject of an entry in a publicly accessible knowledge base, may be reviewed or critiqued by a news organization, may have a dedicated fan website, and may be the subject of discussions on social media. For an owner or creator of such an entity, it may be desirable or even necessary to quickly identify and evaluate the descriptive metadata with which the entity is being tagged by the various sources of news and commentary. The advantages of doing so include enriching the metadata tags already associated with the entity by the entity owner or creator with accurate or laudatory metadata generated by other sources, as well as the prompt correction or removal of metadata tags that are inaccurate or improperly disparaging. (2) Due to the proliferation of potential sources of descriptive commentary made possible by Internet based communications and the growing diversity of social media platforms, timely manual identification and review of metadata tags by human editors in order to match entities on different sources is impracticable. In response, automated solutions for searching out metadata tags for use in performing entity matching have been developed. While offering efficiency advantages over manual tag searching, automated systems are more pro

Claims

1. A metadata aggregation system comprising: a computing platform including a hardware processor and a system memory; a software code stored in the system memory, the software code including a trained entity matching predictive model trained using training data obtained from a reference database; the hardware processor configured to execute the software code to: obtain a plurality of metadata inputs from a plurality of sources; conform the plurality of metadata inputs to a common format; match, using the trained entity matching predictive model, at least some of the conformed plurality of metadata inputs to a same entity to generate a plurality of matches; determine, using the trained entity matching predictive model, a confidence score for each of the plurality of matches; send a confirmation request to at least one human editor for confirmation of each of the plurality of matches having a respective confidence score greater than a first threshold score and less than a second threshold score; and update the reference database, in response to receiving a confirmation of at least one of the plurality of matches as a confirmed match from the at least one human editor, to include the confirmed match in the reference database. || 11. A method for use by a metadata aggregation system including a computing platform having a hardware processor and a system memory storing a software code, the software code including a trained entity matching predictive model trained using training data obtained from a reference database, the method comprising: obtaining, by the software code executed by the hardware processor, a plurality of metadata inputs from a plurality of sources; conforming, by the software code executed by the hardware processor, the plurality of metadata inputs to a common format; matching, by the software code executed by the hardware processor and using the trained entity matching predictive model, at least some of the conformed plurality of metadata inputs to a same entity to generate a plurality of matches; determining, by the software code executed by the hardware processor and using the trained entity matching predictive model, a confidence score for each of the plurality of matches; sending, by the software code executed by the hardware processor, a confirmation request to at least one human editor for confirmation of each of the plurality of matches having a respective confidence score greater than a first threshold score and less than a second threshold score; and updating the reference database, by the software code executed by the hardware processor in response to receiving a confirmation of at least one of the plurality of matches as a confirmed match from the at least one human editor, to include the confirmed match in the reference database.