Techniques are disclosed for characterizing audience engagement with one or more characters in a media content item. In some embodiments, an audience engagement characterization application processes sensor data; such as video data capturing the faces of one or more audience members consuming a media content item, to generate an audience emotion signal. The characterization application also processes the media content item to generate a character emotion signal associated with one or more characters in the media content item. Then, the characterization application determines an audience engagement score based on an amount of alignment and/or misalignment between the audience emotion signal and the character emotion signal.
BACKGROUND Technical Field (1) Embodiments of the present disclosure relate generally to computer science and machine learning and, more specifically, to techniques for characterizing audience engagement based on emotional alignment with characters. Description of the Related Art (2) Producing media content items, such as movies and episodic shows, is oftentimes risky and expensive. Predictions of audience reactions to media content items can inform decisions on whether to produce those media content items. (3) One conventional approach for predicting audience reactions involves showing a media content item to a sample audience that provides feedback on the media content item. For example, a production company might show a pilot episode that is representative of an episodic show to a focus group of volunteers and elicit, from each volunteer, feedback on his or her evaluation of the pilot episode and intent to watch the episodic show. Typically, each volunteer provides feedback via a standardized survey. In some cases, dial testing is also employed. During dial testing, each volunteer turns a knob on a handheld device to provide a real-time signal of his or her opinions towards a media content item. (4) One drawback of the above approaches to predicting audience reactions to a media content item is that these approaches can be susceptible to self-reporting bias. In that regard, volunteers who consume a media content item are required to use their judgment to provide feedback o
1. A computer-implemented method for characterizing engagement with at least one character in a media content item, the method comprising: processing sensor data associated with at least one individual to generate a first signal that indicates one or more emotions expressed by the at least one individual while consuming the media content item or a live event recorded in the media content item; processing the media content item to generate a second signal that indicates one or more emotions expressed by the at least one character in the media content item; and computing a score based on the first signal and the second signal, wherein the score is indicative of at least one of an amount of alignment or an amount of misalignment between the first signal and the second signal, and the computing comprises determining at least one of: 1) whether values of the second signal are predictive of values of the first signal, or 2) whether values of the first signal are predictive of values of the second signal. ||
10. One or more non-transitory computer-readable storage media including instructions that, when executed by at least one processor, cause the at least one processor to perform steps for characterizing engagement with at least one character in a media content item, the steps comprising: processing sensor data associated with at least one individual to generate a first signal that indicates one or more emotions expressed by the at least one individual while consuming the media content item or a live event recorded in the media content item; processing the media content item to generate a second signal that indicates one or more emotions expressed by the at least one character in the media content item; and computing a score based on the first signal and the second signal, wherein the score is indicative of at least one of an amount of alignment or an amount of misalignment between the first signal and the second signal, and the computing comprises determining at least one of: 1) whether values of the second signal are predictive of values of the first signal, or 2) whether values of the first signal are predictive of values of the second signal. ||
19. A system, comprising: one or more sensors that acquire sensor data associated with at least one individual as the at least one individual consumes a media content item or a live event recorded in the media content item; one or more memories storing instructions; and one or more processors that are coupled to the one or more memories and, when executing the instructions, are configured to: process the sensor data to generate a first signal that indicates one or more emotions expressed by the at least one individual, process the media content item to generate a second signal that indicates one or more emotions expressed by at least one character in the media content item, and compute a score based on the first signal and the second signal, wherein the score is indicative of at least one of an amount of alignment or an amount of misalignment between the first signal and the second signal, and the computing comprises determining at least one of: 1) whether values of the second signal are predictive of values of the first signal, or 2) whether values of the first signal are predictive of values of the second signal.