- Number
- 12333258
- Published
- 2025-06-17
- Filed
- 2022-08-24
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Tiwari; Sanchita et al.
- CPC
- G06F40/35; G06F40/30; G10L13/10; G10L25/63; G06N20/00; G06F40/289; G06F40/216; G06F40/284; G10L13/08
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Emotionally enhanced dialogue technique for interactive characters (granted, dup cluster).
Abstract
A system for emotionally enhancing dialogue includes a computing platform having processing hardware and a system memory storing a software code including a predictive model. The processing hardware is configured to execute the software code to receive dialogue data identifying an utterance for use by a digital character in a conversation, analyze, using the dialogue data, an emotionality of the utterance at multiple structural levels of the utterance, and supplement the utterance with one or more emotional attributions, using the predictive model and the emotionality of the utterance at the multiple structural levels, to provide one or more candidate emotionally enhanced utterance(s). The processing hardware further executes the software code to perform an audio validation of the candidate emotionally enhanced utterance(s) to provide a validated emotionally enhanced utterance, and output an emotionally attributed dialogue data providing the validated emotionally enhanced utterance for use by the digital character in the conversation.
Background
BACKGROUND (1) Advances in artificial intelligence (AI) have led to the generation of a variety of digital characters, such as avatars for example, that simulate social interaction. However, conventionally generated AI digital characters typically project a single synthesized persona that tends to lack a distinctive personality and is unable to credibly express emotions. (2) In contrast to conventional interaction simulations by AI digital characters, natural interactions between human beings are more nuanced, varied, and dynamic. For example, conversations between humans are typically expressive of the emotional states of the dialogue partners. That is to say, typical shortcomings of AI digital character interactions include their failure to inflect the words they utter with emotional states such as excitement, disappointment, anxiety, and optimism, to name a few. Thus, there is a need in the art for a dialogue enhancement solution capable of producing emotionally expressive utterances for execution in real-time during a dialogue between a digital character and a dialogue partner such as a human user.
Claims
1. A system comprising: a computing platform having a speech synthesizer, a processing hardware and a system memory storing a software code, the software code including at least one of a trained machine learning (ML) model configured to function as an autoregressive generator or a stochastic model trained using unsupervised or semi-supervised learning; the processing hardware configured to execute the software code to: receive dialogue data identifying an utterance for use by a digital character in a conversation; analyze, using the dialogue data, an emotionality of the utterance at a plurality of structural levels of the utterance; supplement the utterance with one or more emotional attributions, using the at least one of the trained ML model or the stochastic model and the emotionality of the utterance at the plurality of structural levels, to provide a plurality of candidate emotionally enhanced utterances; perform an audio validation of at least some of the plurality of candidate emotionally enhanced utterances to provide a validated emotionally enhanced utterance including a non-verbal vocalization, wherein the audio validation identifies the validated emotionally enhanced utterance as having a best audio quality of the at least some of the plurality of candidate emotionally enhanced utterances; output an emotionally attributed dialogue data providing the validated emotionally enhanced utterance for use by the digital character in the conversation; and synthesize, using the speech synthesizer and the emotionally attributed dialogue data, the validated emotionally enhanced utterance to generate a synthesized speech for utterance by the digital character. ||
12. A method for use by a system including a computing platform having a speech synthesizer, a processing hardware and a system memory storing a software code, the software code including at least one of a trained machine learning (ML) model configured to function as an autoregressive generator or a stochastic model trained using unsupervised or semi-supervised learning: receiving, by the software code executed by the processing hardware, dialogue data identifying an utterance for use by a digital character in a conversation; analyzing, by the software code executed by the processing hardware and using the dialogue data, an emotionality of the utterance at a plurality of structural levels of the utterance; supplementing the utterance with one or more emotional attributions, by the software code executed by the processing hardware and using the at least one of the trained ML model or the stochastic model and the emotionality of the utterance at the plurality of structural levels, to provide a plurality of candidate emotionally enhanced utterances; performing, by the software code executed by the processing hardware, an audio validation of at least some of the plurality of candidate emotionally enhanced utterances to provide a validated emotionally enhanced utterance including a non-verbal vocalization, wherein the audio validation identifies the validated emotionally enhanced utterance as having a best audio quality of the at least some of the plurality of candidate emotionally enhanced utterances; outputting, by the software code executed by the processing hardware, an emotionally attributed dialogue data providing the validated emotionally enhanced utterance for use by the digital character in the conversation; and synthesizing, by the software code executed by the processing hardware using the speech synthesizer and the emotionally attributed dialogue data, the validated emotionally enhanced utterance to generate a synthesized speech for utterance by the digital character.