Outer Rim Archives
Archives · 2022 · 11423879

Granted patent

Verbal cues for high-speed control of a voice-enabled device

Number
11423879
Published
2022-08-23
Filed
2017-07-18
Assignee
Disney Enterprises, Inc.
Inventors
Zajac, III; William Valentine
CPC
A63F13/424; G06F3/167; A63F13/215; G10L15/02
Verdict
Low Notable software
Source
Google Patents · FreePatentsOnline

The keeper's note

Voice-command control technique for interactive devices.

Abstract

A technique for controlling a voice-enabled device using voice commands includes receiving an audio signal that is generated in response to a verbal utterance, generating a verbal utterance indicator for the verbal utterance based on the audio signal, selecting a first command for a voice-controlled application residing within the voice-enabled device based on the verbal utterance indicator, and transmitting the first command to the voice-controlled application as an input.

Background

BACKGROUND OF THE INVENTION Field of the Invention (1) The present disclosure relates generally to computer science, and, more specifically, to verbal cues for high-speed control of a software application. Description of the Related Art (2) Computing devices, such as home automation systems, smart speakers, and gaming consoles, are now equipped with microphones, powerful processors, and advanced speech recognition algorithms. As a result, voice-enabled software applications have come into widespread use. These applications are configured to perform tasks based on voice commands, thereby circumventing the need for a user to provide manual input via a button, control knob, touchscreen, keyboard, mouse, or other input device. For example, using voice commands in conjunction with a voice-enabled software application, a user can modify an audio output volume of a device, select a song to be played by a smart speaker, control a voice-enabled home appliance, etc. Thus, devices configured with voice-enabled software applications (referred to herein as “voice-enabled devices”) are well-suited for situations where the user is unable to perform a manual input, or the use of a manual input device is inconvenient. (3) Despite the positive aspects of voice-enabled devices, trying to control devices using voice input has certain drawbacks. Specifically, the use of complete words or phrases to provide input to a voice-enabled device can be much slower than providing input through other means

Claims

1. A method for controlling a voice-enabled device using voice commands, the method comprising: receiving an audio signal that is generated in response to a verbal utterance; generating a verbal utterance indicator for the verbal utterance based on the audio signal; selecting a first application included in a plurality of applications currently executing on the voice-enabled device that is to receive one or more commands; selecting a first command to transmit to the first application based on the verbal utterance indicator comprising: upon determining that the verbal utterance indicator comprises a phonetic fragment, performing the steps of: performing a first mapping from the verbal utterance indicator to a first complete word via a first mapping table associated with the first application; and performing a second mapping from the first complete word to the first command via a second mapping table associated with the first application; and upon determining that the verbal utterance indicator comprises the first complete word, performing the second mapping from the first complete word to the first command via the second mapping table without performing the first mapping from the verbal utterance indicator to the first complete word via the first mapping table; and transmitting the first command to the first application as an input. || 11. A non-transitory computer-readable storage medium including instructions that, when executed by a processor, cause the processor to perform the steps of: receiving an audio signal that is generated in response to a verbal utterance; generating a verbal utterance indicator for the verbal utterance based on the audio signal; selecting a first application included in a plurality of applications currently executing on the voice-enabled device that is to receive one or more commands; selecting a first command to transmit to the first application based on the verbal utterance indicator comprising: upon determining that the verbal utterance indicator comprises a phonetic fragment, performing the steps of: performing a first mapping from the verbal utterance indicator to a first complete word via a first mapping table associated with the first application; and performing a second mapping from the first complete word to the first command via a second mapping table associated with the first application; and upon determining that the verbal utterance indicator comprises the first complete word, performing the second mapping from the first complete word to the first command via the second mapping table without performing the first mapping from the verbal utterance indicator to the first complete word via the first mapping table; and transmitting the first command to the first application as an input. || 17. A system, comprising: a microphone; a memory storing a speech recognition application and a verbal utterance translator; and one or more processors that are coupled to the memory and the microphone, and when executing the speech recognition application or the verbal utterance translator, are configured to: receive an audio signal that is generated in response to a verbal utterance; generate a verbal utterance indicator for the verbal utterance based on the audio signal; select a first application included in a plurality of applications currently executing on the voice-enabled device that is to receive one or more commands; select a first command to transmit to the first application based on the verbal utterance indicator comprising: upon determining that the verbal utterance indicator comprises a phonetic fragment, performing the steps of: performing a first mapping from the verbal utterance indicator to a first complete word via a first mapping table associated with the first application; and performing a second mapping from the first complete word to the first command via a second mapping table associated with the first application; and upon determining that the verbal utterance indicator comprises the first complete word, performing the second mapping from the first complete word to the first command via the second mapping table without performing the first mapping from the verbal utterance indicator to the first complete word via the first mapping table; and transmit the first command to the first application as an input.