- Number
- 11989922
- Published
- 2024-05-21
- Filed
- 2022-02-18
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Farre Guiu; Miquel Angel et al.
- CPC
- G06N3/08; G06V10/25; G06V10/462; G06N20/20; G06N20/00; G06V20/70
- Verdict
- Set aside image indexing, business
- Source
- Google Patents · FreePatentsOnline
Abstract
A system includes a computing platform having processing hardware, and a memory storing software code. The processing hardware is configured to execute the software code to receive an image having a plurality of image regions, determine a boundary of each of the image regions to identify a plurality of bounded image regions, and identify, within each of the bounded image regions, one or more image sub-regions to identify a plurality of image sub-regions. The processing hardware is further configured to execute the software code to identify, within each of the bounded image regions, one or more first features, respectively, identify, within each of the image sub-regions, one or more second features, respectively, and provided an annotated image by annotating each of the bounded image regions using the respective first features and annotating each of the image sub-regions using the respective second features.
Background
BACKGROUND (1) Due to its nearly universal popularity as a content medium, ever more visual media content is being produced and made available to consumers. As a result, the efficiency with which visual images can be analyzed and rendered searchable has become increasingly important to the producers, owners, and distributors of that visual media content. (2) Annotation and indexing of visual media content for search is typically performed manually by human editors. However, such manual processing is a labor intensive and time consuming process. Moreover, in a typical visual media production environment there may be such a large number of images to be analyzed and indexed that manual processing of those images becomes impracticable. In response, various automated systems for performing image analysis have been developed. While offering efficiency advantages over traditional manual techniques, such automated systems are especially challenged by particular types of visual media content. For example, comics, graphic novels, and Japanese manga present stories about characters with features depicted from the perspectives of drawing artists with different styles that often change over time in different comic or manga issues, within the same comic or manga issue, in different graphic novels in a series, or within the same graphic novel. Moreover, a drawing artist might use different drawing qualities to emphasize different features across the arc of a single storyline. Those conditio
Claims
1. A system comprising: a computing platform having a processing hardware and a system memory storing a software code; the processing hardware configured to execute the software code to: receive an image having a plurality of panels, each of the plurality of panels depicting a plurality of features; determine a respective boundary of each of the plurality of panels to identify a plurality of bounded image regions; identify, within each of the plurality of bounded image regions, respective one or more image sub-regions to identify a plurality of image sub-regions; identify, within each of the plurality of bounded image regions, one or more first features of the plurality of features, respectively; identify, within each of the plurality of image sub-regions, one or more second features of the plurality of features, respectively; and provide an annotated image by annotating each of the plurality of bounded image regions using the respective first features and annotating each of the plurality of image sub-regions using the respective second features. ||
11. A method for use by a system including a computing platform having a processing hardware, and a system memory storing a software code, the method comprising: receiving, by the software code executed by the processing hardware, an image having a plurality of panels, each of the plurality of panels depicting a plurality of features; determining, by the software code executed by the processing hardware, a respective boundary of each of the plurality of panels to identify a plurality of bounded image regions; identifying, by the software code executed by the processing hardware within each of the plurality of bounded image regions, respective one or more image sub-regions to identify a plurality of image sub-regions; identifying, by the software code executed by the processing hardware within each of the plurality of bounded image regions, one or more first features of the plurality of features, respectively; identifying, by the software code executed by the processing hardware within each of the plurality of image sub-regions, one or more second features of the plurality of features, respectively; and providing an annotated image, by the software code executed by the processing hardware, by annotating each of the plurality of bounded image regions using the respective first features and annotating each of the plurality of image sub-regions using the respective second features. ||
20. A system comprising: a processing hardware; a system memory storing a software code; and a plurality of trained machine learning (ML) models; the processing hardware configured to execute the software code to: receive an image having a plurality of image regions; determine, using a first trained ML model of the plurality of trained ML models, a respective boundary of each of the plurality of image regions to identify a plurality of bounded image regions; identify, within each of the plurality of bounded image regions, respective one or more image sub-regions to identify a plurality of image sub-regions; identify within each of the plurality of bounded image regions, using one or more second trained ML model(s) of the plurality of trained ML models, one or more first features, respectively; identify within each of the plurality of image sub-regions, using one or more third trained ML model(s) of the plurality of trained ML models, one or more second features, respectively; and provide an annotated image by annotating each of the plurality of bounded image regions using the respective first features and annotating each of the plurality of image sub-regions using the respective second features.