Application (pre-grant publication)
TECHNIQUES FOR PERFORMING CONTEXTUAL PHRASE GROUNDING
- Number
- 20200272695
- Published
- 2020-08-27
- Filed
- 2019-02-25
- Assignee
- DISNEY ENTERPRISES, INC.
- Inventors
- DOGAN; Pelin, SIGAL; Leonid, GROSS; Markus
- CPC
- G06F40/169; G06F40/30; G06N3/044; G06N3/0442; G06N3/045; G06N3/0464; G06N3/08; G06N3/09
- Verdict
- Set aside NLP phrase grounding, generic language tooling
- Source
- Google Patents · FreePatentsOnline
Abstract
In various embodiments, a phrase grounding model automatically performs phrase grounding for a source sentence and a source image. The phrase grounding model determines that a first phase included in the source sentence matches a first region of the source image based on the first phrase and at least a second phrase included in the source sentence. The phrase grounding model then generates a matched pair that specifies the first phrase and the first region. Subsequently, one or more annotation operations are performed on the source image based on the matched pair. Advantageously, the accuracy of the phrase grounding model is increased relative to prior art solutions where the interrelationships between phrases are typically disregarded.
Background
BACKGROUNDField of the Various Embodiments
Embodiments of the present invention relate generally to natural language and image processing and, more specifically, to techniques for performing contextual phrase grounding.Description of the Related Art
Phrase grounding is the process of matching phrases included in a source sentence to corresponding regions in a source image described by the source sentence. For example, phrase grounding could match the phrases “a small kid,” “blond hair,” and “a cat” included in the source sentence “a small kid with blond hair is kissing a cat” to three different regions in an associated picture or source image. Some examples of high-level tasks that involve phase grounding include image retrieval, image captioning, and visual question answering (i.e., generating natural language answers to natural language questions about an image). For many high-level tasks, manually performing phase grounding is prohibitively time consuming. Consequently, phase grounding techniques are oftentimes performed automatically.
In one approach to automatically performing phase grounding, a noun-matching application breaks a source sentence into constituent noun phrases. The noun-matching application also performs object detection operations that generate bounding boxes corresponding to different objects in a source image. For each noun phrase, the noun-matching application then performs machine-learning operations that independently map each noun