Outer Rim Archives
Archives · 2025 · 20250356839

Application (pre-grant publication)

Artificial Intelligence Based Character-Specific Speech Generation

Number
20250356839
Published
2025-11-20
Filed
2024-05-20
Assignee
Disney Enterprises, Inc.
Inventors
Khalfan; Alif et al.
CPC
G10L13/027; G10L13/047; G06V40/174; G10L15/18; G10L15/183
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

AI character-specific speech-generation technique.

Abstract

A system includes a hardware processor and a memory storing software code, a character database, a language model and an artificial intelligence (AI) model trained to emulate speech by a character. The software code is executed to receive interaction data including a description of speech by a human to a performer impersonating the character and a description of a facial expression by the performer in response, obtain, from the character database, one or more communication trait(s) of the character, and generate, by the language model using the description of the speech and the communication trait(s) as inputs, a character-specific response to the speech. The software code is further executed to synthesize, by the AI model using the character-specific response and the description of the facial expression as inputs, audio data of the character-specific response in a voice of the character, and output the audio data for use by the performer.

Background

BACKGROUND

Performers impersonating famous characters, such as well-known cartoon characters associated with distinctive voices and/or distinctive communication traits for example, may be precluded from speaking using their own voices while performing to avoid inconsistency, incongruity and brand dilution. As a result, a performer impersonating a famous character may be limited to using poses, gestures and physical antics to essentially mime communication in response to a human attempting to interact with the character. Although in some cases that performance may be accompanied by pre-recorded speech by the character in a brand-approved voice and using brand-approved language, the resulting interaction would typically be perceived by the human as lacking spontaneity and immersiveness due to the absence of genuine dialogue. Consequently, there is a need in the art for an automated solution for dynamically generating character-specific speech that is responsive to the emotions and language of a human attempting to engage in dialogue with the character.

Claims

1. A system comprising: a computing platform including a hardware processor and a system memory; the system memory storing a software code, a character database, a language model and an artificial intelligence (AI) model trained to emulate speech by a character; the hardware processor configured to execute the software code to: receive interaction data, the interaction data including a description of a speech by a human to a performer impersonating the character, and a description of a facial expression by the performer in response to the speech; obtain, from the character database, one or more communication traits of the character; generate, by the language model using the description of the speech and the one or more communication traits as inputs, a character-specific response to the speech; synthesize, by the AI model using the character-specific response and the description of the facial expression by the performer as inputs, audio data of the character-specific response in a voice of the character; and output the audio data for use by the performer. || 11. A method for use by a system including a hardware processor and a system memory, the system memory storing a software code, a character database, a language model and an artificial intelligence (AI) model trained to emulate speech by a character, the method comprising: receiving, by the software code executed by the hardware processor, interaction data, the interaction data including a description of a speech by a human to a performer impersonating the character, and a description of a facial expression by the performer in response to the speech; obtaining from the character database, by the software code executed by the hardware processor, one or more communication traits of the character; generating, by the language model using the description of the speech and the one or more communication traits as inputs, a character-specific response to the speech; synthesizing, by the AI model using the character-specific response and the description of the facial expression by the performer as inputs, audio data of the character-specific response in a voice of the character; and outputting, by the software code executed by the hardware processor, the audio data for use by the performer.