Embodiments provide techniques for distributing supplemental content based on content entities within video content. Embodiments include analyzing video data to identify a known content entity within two or more frames of the video data. For each of the two or more frames, a region of pixels within the respective frame is determined that corresponds to the known content entity. Embodiments further include determining supplemental content corresponding to the known content entity. A watermark is embedded at a first position within the video data, such that the watermark corresponds to an identifier associated with the determined supplemental content. Upon receiving a message specifying the identifier, embodiments include transmitting the supplement content to a client device for output together with the video data.
BACKGROUNDField of the Invention(1) The present disclosure relates to providing media content, and more specifically, to techniques for embedding data within video content for use in retrieving supplemental content corresponding to a content entity depicted within the video content.Description of the Related Art(2) A number of different techniques exist today for delivering video content to users. Generally speaking, existing systems for delivering video content, such as over-the-air broadcasts, cable television service, Internet television service, telephone network television service, satellite television service, satellite radio service, websites, etc., provide a relatively impersonalized, generic experience to all viewers. For example, with respect to broadcast television, all viewers of a given television network station receive essentially the same content in essentially the same order.(3) In addition to providing the video content to users, content providers can also provide supplemental content that corresponds to content entities shown in the video content. For instance, a particular scene of the video content could show a particular actor playing a particular fictional character, and supplemental content about the particular fictional character could be provided along with the video content. For example, such supplemental content could include concept art and a biography for the fictional character. Moreover, such supplemental content can be shown for other content
1. A computer-implemented method comprising: analyzing video data for a first instance of video content to identify a character in a scene depicted within two or more frames of a plurality of frames of the video data, wherein the character is identified in at least one frame of the two or more frames based at least in part on a corresponding identification in at least one adjacent frame; determining, for each of the two or more frames, a region of pixels within the respective frame that correspond to the character, including normalizing the region of pixels across the two or more frames; generating a mapping between the character, and a data value pair specifying (i) a content identifier that uniquely identifies the first instance of video content and (ii) a timestamp corresponding to a first position within the first instance of video content where the character was identified; receiving, from a remote client device, a message specifying (i) the content identifier that uniquely identifies the first instance of video content and (ii) the timestamp corresponding to the first position within the first instance of video content; determining that the message pertains to the character by accessing the mapping using the timestamp and the content identifier specified within the message; determining supplemental content having a predefined correspondence with the character; and transmitting at least a portion of the supplemental content having the predefined correspondence with the character to the remote client device, for output together with the video data.
16. A non-transitory computer-readable medium containing program code that, when executed by operation of one or more computer processors, performs an operation comprising: analyzing video data for a first instance of video content to identify a character in a scene depicted within two or more frames of a plurality of frames of the video data, wherein the character is identified in at least one frame of the two or more frames based at least in part on a corresponding identification in at least one adjacent frame; determining, for each of the two or more frames, a region of pixels within the respective frame that correspond to the character, including normalizing the region of pixels across the two or more frames; generating a mapping between the character and a data value pair specifying (i) a content identifier that uniquely identifies the first instance of video content and (ii) a timestamp corresponding to a first position within the first instance of video content where the character was identified; receiving, from a remote client device, a message specifying (i) a content identifier that uniquely identifies the first instance of video content and (ii) the timestamp corresponding to the first position within the first instance of video content; determining that the message pertains to the character by accessing the mapping using the timestamp and the content identifier specified within the message; determining supplemental content having a predefined correspondence with the character; and transmitting at least a portion of the supplemental content having the predefined correspondence with the character to the remote client device, for output together with the video data.
20. A system comprising: one or more computer processors; and a memory containing a program that, when executed by the one or more computer processors, performs an operation comprising: analyzing video data for a first instance of video content to identify a character in a scene depicted within two or more frames of a plurality of frames of the video data, wherein the character is identified in at least one frame of the two or more frames based at least in part on a corresponding identification in at least one adjacent frame; determining, for each of the two or more frames, a region of pixels within the respective frame that correspond to the character, including normalizing the region of pixels across the two or more frames; generating a mapping between the character and a data value pair specifying (i) a content identifier that uniquely identifies the first instance of video content and (ii) a timestamp corresponding to a first position within the first instance of video content where the character was identified; receiving, from a remote client device, a message specifying (i) a content identifier that uniquely identifies the first instance of video content and (ii) the timestamp corresponding to the first position within the first instance of video content; determining that the message pertains to the character by accessing the mapping using the timestamp and the content identifier specified within the received message; determining supplemental content having a predefined correspondence with the character; and transmitting at least a portion of the supplemental content having the predefined correspondence with the character to the remote client device, for output together with the video data.