Application (pre-grant publication)
TECHNIQUES FOR INTERPRETING SPOKEN INPUT USING NON-VERBAL CUES
- Number
- 20210104241
- Published
- 2021-04-08
- Filed
- 2019-10-04
- Assignee
- DISNEY ENTERPRISES, INC.
- Inventors
- DOGGETT; Erika, NOCON; Nathan, MODI; Ashutosh, SENGIR; Joseph Charles, MCCOY; Maxwell
- CPC
- G10L15/24; G10L15/063; G10L15/22; G10L15/26; G06F3/167; G10L15/25
- Verdict
- Set aside generic speech-input interpretation research
- Source
- Google Patents · FreePatentsOnline
Abstract
In various embodiments, a communication fusion application enables other software application(s) to interpret spoken user input. In operation, a communication fusion application determines that a prediction is relevant to a text input derived from a spoken input received from a user. Subsequently, the communication fusion application generates a predicted context based on the prediction. The communication fusion application then transmits the predicted context and the text input to the other software application(s). The other software application(s) perform additional action(s) based on the text input and the predicted context. Advantageously, by providing additional, relevant information to the software application(s), the communication fusion application increases the level of understanding during interactions with the user and the overall user experience is improved.
Background
BACKGROUND Field of the Various Embodiments
The various embodiments relate generally to computer science and natural language understanding and, more specifically, to techniques for interpreting spoken input using non-verbal cues. Description of the Related Art
To make user interactions with machine-based systems seem more natural to users, some service providers implement chat-based applications within user devices that are designed to allow the devices to communicate verbally with users. Such chat-based applications are commonly referred to as a “chatbots.” In a typical user interaction, a speech-to-text model translates spoken input to text input that can include any number of words. The speech-to-text model transmits the text input to the chat-based application, and, in response, the chat-based application generates text output based on the text input. The chat-based application then configures a speech synthesizer to translate the text output to spoken output, which is then transmitted from the device to the user.
One drawback of using chat-based applications is that speech-to-text models usually do not take non-verbal cues into account when translating spoken input to text input. Consequently, in many use cases, chat-based applications are not able to respond properly to user input. More specifically, when speaking, humans oftentimes communicate additional information along with spoken words that impacts how the spoken words should be interpreted. T