Granted patent
Metadata aggregation using a trained entity matching predictive model
- Number
- 11314706
- Published
- 2022-04-26
- Filed
- 2020-08-21
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Stoafer; Christopher C., Pujol; Jordi Badia, Bravo; Francesc Josep Guitart, Martin; Marc Junyent, Farre Guiu; Miquel Angel, Lawson; Calvin, Luerken; Erick L.
- CPC
- G06F11/3419; G06F11/3433; G06F16/1794; G06F16/215; G06F16/217; G06F16/24542; G06F16/24549; G06F16/2462; G06F16/3346; G06F16/38; G06F16/383; G06F17/11; G06F17/18; G06F30/27; G06F9/4881; G06N7/01
- Verdict
- Set aside metadata aggregation, data-engineering business
- Source
- Google Patents · FreePatentsOnline
Abstract
A metadata aggregation system includes a computing platform having a hardware processor and a memory storing a software code including a trained entity matching predictive model trained using training data obtained from a reference database. The hardware processor executes the software code to obtain metadata inputs from multiple sources, conform the metadata inputs to a common format, match, using the trained entity matching predictive model, at least some of the conformed metadata inputs to the same entity, and determine, using the trained entity matching predictive model, a confidence score for each match. The software code further sends a request to one or more human editor(s) for confirmation of each match having a confidence score greater than a first threshold and less than a second threshold, and updates the reference database, in response to receiving a confirmation that at least one match is a confirmed match, to include the confirmed match.
Background
BACKGROUND (1) Popular movies, television programs, sports teams, and other pop culture “entities” are typically the subjects of extensive commentary from a variety of different sources. For example, a movie or movie franchise may be the subject of an entry in a publicly accessible knowledge base, may be reviewed or critiqued by a news organization, may have a dedicated fan website, and may be the subject of discussions on social media. For an owner or creator of such an entity, it may be desirable or even necessary to quickly identify and evaluate the descriptive metadata with which the entity is being tagged by the various sources of news and commentary. The advantages of doing so include enriching the metadata tags already associated with the entity by the entity owner or creator with accurate or laudatory metadata generated by other sources, as well as the prompt correction or removal of metadata tags that are inaccurate or improperly disparaging. (2) Due to the proliferation of potential sources of descriptive commentary made possible by Internet based communications and the growing diversity of social media platforms, timely manual identification and review of metadata tags by human editors in order to match entities on different sources is impracticable. In response, automated solutions for searching out metadata tags for use in performing entity matching have been developed. While offering efficiency advantages over manual tag searching, automated systems are more pro