Outer Rim Archives
Archives · 2020 · 20200272695

Application (pre-grant publication)

TECHNIQUES FOR PERFORMING CONTEXTUAL PHRASE GROUNDING

Number
20200272695
Published
2020-08-27
Filed
2019-02-25
Assignee
DISNEY ENTERPRISES, INC.
Inventors
DOGAN; Pelin, SIGAL; Leonid, GROSS; Markus
CPC
G06F40/169; G06F40/30; G06N3/044; G06N3/0442; G06N3/045; G06N3/0464; G06N3/08; G06N3/09
Verdict
Set aside NLP phrase grounding, generic language tooling
Source
Google Patents · FreePatentsOnline

Abstract

In various embodiments, a phrase grounding model automatically performs phrase grounding for a source sentence and a source image. The phrase grounding model determines that a first phase included in the source sentence matches a first region of the source image based on the first phrase and at least a second phrase included in the source sentence. The phrase grounding model then generates a matched pair that specifies the first phrase and the first region. Subsequently, one or more annotation operations are performed on the source image based on the matched pair. Advantageously, the accuracy of the phrase grounding model is increased relative to prior art solutions where the interrelationships between phrases are typically disregarded.

Background

BACKGROUNDField of the Various Embodiments

Embodiments of the present invention relate generally to natural language and image processing and, more specifically, to techniques for performing contextual phrase grounding.Description of the Related Art

Phrase grounding is the process of matching phrases included in a source sentence to corresponding regions in a source image described by the source sentence. For example, phrase grounding could match the phrases “a small kid,” “blond hair,” and “a cat” included in the source sentence “a small kid with blond hair is kissing a cat” to three different regions in an associated picture or source image. Some examples of high-level tasks that involve phase grounding include image retrieval, image captioning, and visual question answering (i.e., generating natural language answers to natural language questions about an image). For many high-level tasks, manually performing phase grounding is prohibitively time consuming. Consequently, phase grounding techniques are oftentimes performed automatically.

In one approach to automatically performing phase grounding, a noun-matching application breaks a source sentence into constituent noun phrases. The noun-matching application also performs object detection operations that generate bounding boxes corresponding to different objects in a source image. For each noun phrase, the noun-matching application then performs machine-learning operations that independently map each noun

Claims

1. A computer-implemented method for performing automated phase grounding operations, the method comprising: determining that a first phase included in a source sentence matches a first region of a source image based on the first phrase and at least a second phrase included in the source sentence; and generating a first matched pair that specifies the first phrase and the first region, wherein one or more annotation operations are subsequently performed on the source image based on the first matched pair. 11. One or more non-transitory computer readable media including instructions that, when executed by one or more processors, cause the one or more processors to perform automated phase grounding operations by performing the steps of: determining that a first phase included in a source sentence matches a first region of a source image based on the first phrase and at least a second phrase included in the source sentence; and generating a first matched pair that specifies the first phrase and the first region, wherein one or more annotation operations are subsequently performed on the source image based on the first matched pair. 20. A system, comprising: one or more memories storing instructions; and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to: determine that a first phase included in a source sentence matches a first region of a source image based on the first phrase and at least a second phrase included in the source sentence; and generate a first matched pair that specifies the first phrase and the first region, wherein one or more annotation operations are subsequently performed on the source image based on the first matched pair.