- Number
- 11064268
- Published
- 2021-07-13
- Filed
- 2018-03-23
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Farre Guiu; Miquel Angel, Petrillo; Matthew C., Alfaro Vendrell; Monica, Junyent Martin; Marc, Ettinger; Katharine S., Binder; Evan A., Accardo; Anthony M., Swerdlow; Avner
- CPC
- G06F16/7834; H04N21/458; H04N21/845; H04N21/8405; G06V20/48; H04N21/8402; G06F16/783; H04N21/8541
- Verdict
- Set aside content metadata mapping, business
- Source
- Google Patents · FreePatentsOnline
Abstract
According to one implementation, a media content annotation system includes a computing platform having a hardware processor and a system memory storing a software code. The hardware processor executes the software code to receive a first version of media content and a second version of the media content altered with respect to the first version, and to map each of multiple segments of the first version of the media content to a corresponding one segment of the second version of the media content. The software code further aligns each of the segments of the first version of the media content with its corresponding one segment of the second version of the media content, and utilizes metadata associated with each of at least some of the segments of the first version of the media content to annotate its corresponding one segment of the second version of the media content.
Background
BACKGROUND (1) Media content, such as movie or television (TV) content for example, is often produced in multiple versions that, while including much of the same core content, may differ in resolution, through the deletion of some original content, or through the addition of advertising or ancillary content. One example of such versioning is a censored version of a movie in which some scenes from the original master version of the movie are removed. Another example of such versioning is a broadcast version of TV programming content in which the content included in the original master version of the TV content is supplemented with advertising content. (2) Despite the evident advantages of versioning media content to accommodate the tastes and sensibilities of a target audience, or the requirements of advertisers sponsoring distribution of the media content, the consistent annotation of media content across its multiple versions has presented significant challenges. Those challenges arise due to the change in temporal location of a particular segment of the media content from one version to another. In the conventional art, the process of mapping metadata from a master version of media content to other versions of that content is a manual process the may require hours of work by a human editor. SUMMARY (3) There are provided systems and methods for performing media content metadata mapping, substantially as shown in and/or described in connection with at least one of the figure
Claims
1. A media content annotation system comprising: a computing platform including a hardware processor and a system memory; a software code stored in the system memory; the hardware processor configured to execute the software code to: receive a first version of a media content and a second version of the media content, wherein the second version of the media content is created by altering the first version of the media content; map each of a plurality of first video shots of the first version of the media content to a corresponding one of a plurality of second video shots of the second version of the media content based on determining shot visual feature similarities between first video shot features of the plurality of first video shots and second video shot features of the corresponding one of the plurality of second video shots; align, based on mapping the plurality of first video shots, each of the plurality of first video shots of the first version of the media content with the corresponding one of the plurality of second video shots of the second version of the media content, wherein each of the plurality of first video shots includes a plurality of first video frames, and each of the plurality of second video shots includes a plurality of second video frames; after aligning each of the plurality of first video shots with the corresponding one of the plurality of second video shots, for each pair of the aligned one of the plurality of first video shots and the corresponding one of the plurality of second video shots: map each of the plurality of first video frames to a corresponding one of the plurality of second video frames based on determining frame visual feature similarities between first video frame features of the plurality of first video frames and second video frame features of the corresponding one of the plurality of second video frames; align, based on mapping the plurality of first video frames, each of the plurality of first video frames with the corresponding one of the plurality of second video frames; and utilize metadata associated with a subset of the plurality of first video shots or a subset of the plurality of first video frames of the first version of the media content, to annotate the respective corresponding one of the plurality of second video shots or the plurality of second video frames of the second version of the media content. ||
6. A method for use by a media content annotation system including a computing platform having a hardware processor and a system memory storing a software code, the method comprising: receiving, using the hardware processor, a first version of a media content and a second version of the media content, wherein the second version of the media content is created by altering the first version of the media content; mapping, using the hardware processor, each of a plurality of first video shots of the first version of the media content to a corresponding one of a plurality of second video shots of the second version of the media content based on determining shot visual feature similarities between first video shot features of the plurality of first video shots and second video shot features of the corresponding one of the plurality of second video shots; aligning, using the hardware processor and based on mapping the plurality of first video shots, each of the plurality of first video shots of the first version of the media content with the corresponding one of the plurality of second video shots of the second version of the media content, wherein each of the plurality of first video shots includes a plurality of first video frames, and each of the plurality of second video shots includes a plurality of second video frames; after aligning each of the plurality of first video shots with the corresponding one of the plurality of second video shots, for each pair of the aligned one of the plurality of first video shots and the corresponding one of the plurality of second video shots: mapping, using the hardware processor, each of the plurality of first video frames to a corresponding one of the plurality of second video frames based on determining frame visual feature similarities between first video frame features of the plurality of first video frames and second video frame features of the corresponding one of the plurality of second video frames; aligning, using the hardware processor and based on mapping the plurality of first video frames, each of the plurality of first video frames with the corresponding one of the plurality of second video frames; and utilizing, using the hardware processor, metadata associated with a subset of the plurality of first video shots or a subset of the plurality of first video frames of the first version of the media content, to annotate the respective corresponding one of the plurality of second video shots or the plurality of second video frames of the second version of the media content.