Outer Rim Archives
Archives · 2017 · 20170061959

Application (pre-grant publication)

Systems and Methods For Detecting Keywords in Multi-Speaker Environments

Number
20170061959
Published
2017-03-02
Filed
2015-09-01
Assignee
Disney Enterprises, Inc.
Inventors
Lehman; Jill Fain et al.
CPC
G10L15/08
Verdict
Set aside keyword detection in multi-speaker audio, generic speech-recognition tooling
Source
Google Patents · FreePatentsOnline

Abstract

There is provided a system for keyword recognition comprising a memory storing a keyword recognition application, a processorexecuting the keyword recognition application to receive a digitized speech from an analog-to-digital (A/D) converter, divide the digitized speech into a plurality of speech segments having a first speech segment, calculate a first probability of distribution of a first keyword in the first speech segment, determine that a first fraction of the first speechsegment includes the first keyword, in response to comparing the first probability of distribution with a first threshold associated with the first keyword, calculate a second probability of distribution of a second keyword in the first speech segment, and determine that a second fraction of the first speech segment includes the second keyword, in response to comparing the second probability of distribution with a second threshold associated with the second keyword.

Background

BRIEF DESCRIPTION OF THE DRAWINGS

FIG. 1 shows a diagram of an exemplary system for detecting keywords in multi-speaker environments, according to one implementation of the present disclosure;

FIG. 2 shows an exemplaryinput speech for processing by the system of FIG. 1, according to one implementation of the present disclosure;

FIG. 3 shows an exemplary speech segment for processing by the system of FIG. 1, according to one implementation of the present disclosure; and

FIG. 4 shows a flowchart illustrating of an exemplary method of detecting keywords in multi-speaker environments, according to one implementation of the present disclosure.DETAILED DESCRIPTION

The following description contains specific information pertaining to implementations in the present disclosure. The drawings in the present application and theiraccompanying detailed description are directed to merely exemplary implementations. Unless noted otherwise, like or corresponding elements among the figures may be indicated by like or corresponding reference numerals. Moreover, the drawings and illustrations in the present application are generally not to scale, and are not intended to correspond to actualrelative dimensions.

FIG. 1 shows a diagram of an exemplary system for detecting keywords in a multi-speaker environment, according to one implementation of the present disclosure. System 100 includes microphone 105, device 110, and peripheral component 195. Device 110 inc

Claims

1. A system for keyword recognition, the system comprising: a microphone configured to receive an input speech; an analog-to-digital (A/D) converter configured to convert the input speech from an analog form to a digital form and generate a digitized speech; a memory storing a keyword recognition application; a hardware processor executing the keyword recognition application to: receive the digitized speech from the A/D converter; divide the digitized speech into a plurality of speech segments having afirst speech segment; calculate a first probability of distribution of a first keyword inthe first speech segment; determine that a first fraction of the first speech segment includes the first keyword, in response to comparing the first probability of distribution with a first threshold associated with the first keyword; calculate a second probability of distribution of a second keyword in the first speech segment; and determine that a second fraction of the first speech segment includes the second keyword, in response to comparingthe second probability of distribution with a second threshold associated with the secondkeyword. 11. A method of keyword recognition, for use with a system having a microphone, an analog-to-digital (A/D) converter, a memory including a keyword recognition application, and a hardware processor, the method comprising: receiving, using the hardware processor, a digitized speech from the A/D converter; dividing, using the hardware processor, the digitized speech into a plurality of speech segments having a first speech segment; calculating, using the hardware processor, a first probability of distribution of a first keywordin the first speech segment; determining, using the hardware processor, that a first fraction of the first speech segment includes the first keyword, in response to comparing the first probability of distribution with a first threshold associated with the first keyword; calculating, using the hardware processor, a second probability of distribution of a second keyword in the first speech segment; and determining, using the hardware processor, that a second fraction of the first speech segment includes the second keyword, in response to comparing the second probability of distribution with a second threshold associated with the second keyword. 20. A system for keyword recognition, the system comprising: a microphone configured to receive an input speech; an analog-to-digital (A/D) converter configured to convert the input speech from an analog form to a digital form and generate a digitized speech; a memory storing a keyword recognition application; a hardware processor executing the keyword recognition application to: receive the digitized speech from the A/D converter; divide the digitized speech into a plurality of speech segments, including a firstspeech segment including a plurality of keywords and background, wherein background includes portions of the first speech segment that do not contain keywords; convert the first speech segment to a feature vector sequence; model a plurality of keyword probability distributions from the feature vector sequence, wherein each keyword probability distribution of the plurality of keyword probability distributions corresponds to a keyword of the plurality of keywords; model a background probability distribution from the feature vector sequence; model the first speech segment as a combination of a plurality of keyword vectors and a plurality of background vectors; model a speech segment probability distribution as a mixture of the plurality of keyword probability distributions and the background probability distribution; estimate a plurality of keyword mixture weights corresponding to the plurality of keyword probability distributions and a background mixture weight corresponding to the background probability distribution using an any maximum-likelihood technique; equate each keyword mixture weight of the plurality of keyword mixture weights to a corresponding plurality of probabilities of each keyword of the plurality of keywords and to a corresponding plurality of fractions of the first speech segment that contain each keyword of the plurality of keywords.