Outer Rim Archives
Archives · 2019 · 10311863

Granted patent

Classifying segments of speech based on acoustic features and context

Number
10311863
Published
2019-06-04
Filed
2016-09-02
Assignee
Disney Enterprises, Inc.
Inventors
Lehman; Jill Fain et al.
CPC
G10L15/22; G10L15/1815
Verdict
Set aside classifying speech segments by acoustic features, generic speech analytics
Source
Google Patents · FreePatentsOnline

Abstract

There is provided a system including a microphone configured to receive an input speech, an analog to digital (A/D) converter configured to convert the input speech to a digital form and generate a digitized speech including a plurality of segments having acoustic features, a memory storing an executable code, and a processor executing the executable code to extract a plurality of acoustic feature vectors from a first segment of the digitized speech, determine, based on the plurality of acoustic feature vectors, a plurality of probability distribution vectors corresponding to the probabilities that the first segment includes each of a first keyword, a second keyword, both the first keyword and the second keyword, a background, and a social speech, and assign a first classification label to the first segment based on an analysis of the plurality of probability distribution vectors of one or more segments preceding the first segment and the probability distribution vectors of the first segment.

Background

BACKGROUND(1) As speech recognition technology has advanced, voice-activated devices have become more and more popular and have found new applications. Today, an increasing number of mobile phones, in-home devices, and automobile devices include speech or voice recognition capabilities. Although the speech recognition modules incorporated into such devices are trained to recognize specific keywords, they tend to be unreliable. This is because the specific keywords may appear in a spoken sentence and be incorrectly recognized as voice commands by the speech recognition module when not intended by the user. Also, in some cases, the specific keywords intended to be taken as commands may not be recognized by the speech recognition module, because the specific keywords may be ignored for appearing in between other spoken words. Of course, both situations can frustrate the user and cause the user to give up or resort to inputting the commands manually, speak the keywords numerous times, or turn off the voice recognition.SUMMARY(2) The present disclosure is directed to classifying segments of speech based on acoustic features and context, substantially as shown in and/or described in connection with at least one of the figures, as set forth more completely in the claims.

Claims

1. A system comprising: a microphone configured to receive an input speech; an analog to digital (A/D) converter configured to convert the input speech from an analog form to a digital form and generate a digitized speech including a plurality of segments having acoustic features; a memory storing an executable code; and a hardware processor executing the executable code to: extract a plurality of acoustic feature vectors from the first segment of the plurality of segments of the digitized speech; determine, based on the plurality of acoustic feature vectors, a plurality of probability distribution vectors corresponding to the probabilities that the first segment of the digitized speech includes each of a first keyword, a second keyword, both the first keyword and the second keyword, a background, and a social speech; assign a first classification label to the first segment based on an analysis of the plurality of probability distribution vectors of one or more segments preceding the first segment and the probability distribution vectors of the first segment, wherein the assigning of the first classification label to the first segment includes distinguishing between the first keyword and an audio event that is one of the social speech and the background; and execute a first action associated with a first command when the first classification label includes the first command. | 9. A method for use with a system having a hardware processor, and a non-transitory memory, the method comprising: extracting, using the hardware processor, a plurality of acoustic feature vectors from a first segment of a plurality of segments of a digitized speech; determining, using the hardware processor and based on the plurality of acoustic feature vectors, a plurality of probability distribution vectors corresponding to the probabilities that first segment of the digitized speech includes each of a first keyword, a second keyword, both the first keyword and the second keyword, a background, and a social speech; assigning, using the hardware processor, a first classification label to the first segment based on the plurality of probability distribution vectors of one or more segments preceding the first segment and the probability distribution vectors of the first segment, wherein the assigning of the first classification label to the first segment includes distinguishing between the first keyword and an audio event that is one of the social speech and the background; and executing a first action associated with a first command when the first classification label includes the first command.