- Number
- 20250182741
- Published
- 2025-06-05
- Filed
- 2023-12-01
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Kumar; Komath Naveen et al.
- CPC
- G10L15/02; G10L13/027; G10L13/033; G10L15/187; G10L15/22; G10L15/1807; G10L13/10
- Verdict
- Low Notable software
- Source
- Google Patents · FreePatentsOnline
The keeper's note
Interactive character-expression speech-rendering technique.
Abstract
A system includes a hardware processor and a memory storing software code and a natural language understanding (NLU) model. The hardware processor executes the software code to receive audio input including speech by a human speaker, produce a text transcription of the audio input, identify, using the NLU model and the text transcription, a segment of interest of the audio input that includes a feature of interest, and analyze one or more audio characteristic(s) of the feature of interest. The software code is further executed to identify, using the text transcription, a text string corresponding to the feature of interest, generate a response to the audio input that includes the text string, and modify the response using the audio characteristic(s) to produce an output response in which the text string is uttered in a characteristic voice of a non-human social agent using a word pronunciation utilized by the human speaker.
Background
BACKGROUND
One of the characteristic features of human interaction is variety of expression, including variations in the way certain words are pronounced. First names, surnames and place names, for example, may be pronounced differently by different people, or may be pronounced differently under different circumstances. For instance, the place name St. John, as well as the surname St. John in American English, is typically pronounced “Saint John.” However, in British English, when used as a first name or other given name, St. John is typically pronounced “Sinjin.”
In order for a non-human social agent, such as one embodied as an artificial intelligence interactive character, for example, to build rapport with a human interacting with the social agent, it is desirable for the social agent to be able to mirror the pronunciations utilized by the human speaker. However, conventional approaches to interpreting human speech and generating responsive expressions for use by a social agent rely on speech-to-text and text-to-speech transcription techniques that can undesirably produce dissonant results. Consequently, there is a need in the art for a solution enabling a non-human social agent to vary its pronunciation to agree with that of a human speaker with whom the social agent interacts.
Claims
1. A system comprising: a computing platform having a hardware processor and a system memory; the system memory storing a software code and a natural language understanding (NLU) model; the hardware processor configured to execute the software code to: receive an audio input, the audio input including speech by a human speaker; produce a text transcription of the audio input; identify, using the NLU model and the text transcription, a segment of interest of the audio input, the segment of interest including a feature of interest; analyze one or more audio characteristics of the feature of interest; identify, using the text transcription, a text string corresponding to the feature of interest; generate a response to the audio input, the response including the text string; and modify the response using the one or more audio characteristics of the feature of interest to produce an output response in which the text string is uttered in a characteristic voice of a non-human social agent using a word pronunciation utilized by the human speaker. ||
8. A method for use by a system including a computing platform having a hardware processor and a system memory, the system memory storing a software code and a natural language understanding model (NLU), the method comprising: receiving, by the software code executed by the hardware processor, an audio input, the audio input including speech by a human speaker; producing, by the software code executed by the hardware processor, a text transcription of the audio input; identifying, by the software code executed by the hardware processor and using the NLU model and the text transcription, a segment of interest of the audio input, the segment of interest including a feature of interest; analyzing, by the software code executed by the hardware processor, one or more audio characteristics of the feature of interest; identifying, by the software code executed by the hardware processor and using the text transcription, a text string corresponding to the feature of interest; generating, by the software code executed by the hardware processor, a response to the audio input, the response including the text string; and modifying the response, by the software code executed by the hardware processor, using the one or more audio characteristics of the feature of interest to produce an output response in which the text string is uttered in a characteristic voice of a non-human social agent using a word pronunciation utilized by the human speaker. ||
15. A computer-readable non-transitory medium having stored thereon instructions, which when executed by a hardware processor, instantiate a method comprising: receiving an audio input, the audio input including speech by a human speaker; producing a text transcription of the audio input; identifying, using an NLU model and the text transcription, a segment of interest of the audio input, the segment of interest including a feature of interest; analyzing one or more audio characteristics of the feature of interest; identifying, using the text transcription, a text string corresponding to the feature of interest; generating a response to the audio input, the response including the text string; and modifying the response, using the one or more audio characteristics of the feature of interest, to produce an output response in which the text string is uttered in a characteristic voice of a non-human social agent using a word pronunciation utilized by the human speaker.