- Number
- 9570069
- Published
- 2017-02-14
- Filed
- 2014-09-09
- Assignee
- Disney Enterprises, Inc.
- Inventors
- Lehman; Jill F. et al.
- CPC
- G10L15/16
- Verdict
- Set aside word-spotting in continuous speech via memory networks, generic speech-recognition tooling
- Source
- Google Patents · FreePatentsOnline
Abstract
Systems, methods, and computer program products to detect a keyword in speech, by generating, from a sequence of spectral feature vectors generated from the speech, a plurality of blocked feature vector sequences, and analyzing, by a neural network, each of the plurality of blocked feature vector sequences to detect the presence of the keyword in the speech.
Background
BRIEF DESCRIPTION OF THE DRAWINGS(1) So that the manner in which the above recited aspects are attained and can be understood in detail, a more particular description of embodiments of the disclosure, briefly summarized above, may be had by reference to the appended drawings.(2) It is to be noted, however, that the appended drawings illustrate only typical embodiments of this disclosure and are therefore not to be considered limiting of its scope, for the disclosure may admit to other equally effective embodiments.(3) FIG. 1 is a flow diagram illustrating techniques for sectioned memory networks to perform online word-spotting in continuous speech, according to one embodiment.(4) FIG. 2 is a block diagram illustrating a system for sectioned memory networks to perform online word-spotting in continuous speech, according to one embodiment.(5) FIG. 3 illustrates components of a neural network,according to one embodiment.(6) FIG. 4 illustrates a method to provide sectioned memory networks to perform online word-spotting in continuous speech, according to one embodiment.(7) FIG. 5 illustrates components of a keyword application, according to one embodiment.DETAILED DESCRIPTION(8) Embodiments disclosed herein provide techniques for identifying keywords in human speech directly through a neural network, without having to search a keywordlattice. Specifically, embodiments disclosed herein use a recurrent neural network architecture to identify words, and not non-word phonemes, such
Claims
1. A method to detect a keyword in speech, comprising: generating, from a sequence of spectral featurevectors generated from the speech, a plurality of blocked feature vector sequences, wherein each of the plurality of blocked feature vector sequences comprises a respective subsetof the plurality of feature vectors, wherein each of the plurality of blocked feature vector sequences overlap adjacent blocked feature vector sequences by a predefined count of feature vector sequences; and analyzing, by a neural network executing on one or more computer processors, each of the plurality of blocked feature vector sequences to detect the presence of the keyword in the speech.
8. A computer program product to detect a keyword in speech, the computer program product comprising: a non-transitory computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code executable by a processor to perform an operation comprising: generating, froma sequence of spectral feature vectors of the speech, a plurality of blocked feature vector sequences, wherein each of the plurality of blocked feature vector sequences comprises a respective subset of the plurality of feature vectors, wherein each of the plurality of blocked feature vector sequences overlap adjacent blocked feature vector sequences by a predefined count of feature vector sequences; and generating, by a neural network, the plurality of blocked feature vector sequences to detect the presence of the keyword in the speech.
14. A system, comprising: a computer processor; and a memory containing a program, which when executed by the computer processor, performs an operation to detect a keyword in speech, the operation comprising: generating, from a sequence of spectral feature vectors generated from the speech, a plurality of blocked feature vector sequences, wherein each ofthe plurality of blocked feature vector sequences comprises a respective subset of the plurality of feature vectors, wherein each of the plurality of blocked feature vector sequences overlap adjacent blocked feature vector sequences by a predefined count of feature vector sequences; and analyzing, by a neural network, each of the plurality of blocked feature vector sequences to detect the presence of the keyword in the speech.
19. A method to detect a keyword in speech, comprising: generating, from a sequence of spectral feature vectors generated from the speech, a plurality of blocked feature vector sequences; generating, by a neural network executing on one or more computer processors and based on each of the plurality of blocked feature vector sequences, a sequence of labels, wherein each label of the sequence of labels indicates whether the keyword is present in a corresponding blocked feature vector sequence; and smoothing the sequence of labels to determine whether thekeyword is in the speech.
20. A computer program product to detect a keyword in speech, the computer program product comprising: a non-transitory computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code executable by a processor to perform an operation comprising: generating, from a sequence of spectral feature vectors generated from the speech, a plurality of blocked feature vector sequences; generating, by a neural network executing on one or more computer processors and based on each of the plurality of blocked feature vector sequences, a sequence of labels, wherein each label of the sequence of labels indicates whether the keyword is present in a corresponding blocked feature vector sequence; and smoothing the sequence of labels to determine whether the keyword is in the speech.
21. A system, comprising: a computerprocessor; and a memory containing a program, which when executed by the computer processor, performs an operation to detect a keyword in speech, the operation comprising: generating, from a sequence of spectral feature vectors generated from the speech, a plurality ofblocked feature vector sequences; generating, by a neural network executing on one or more computer processors and based on each of the plurality of blocked feature vector sequences, a sequence of labels, wherein each label of the sequence of labels indicates whether the keyword is present in a corresponding blocked feature vector sequence; and smoothing the sequence of labels to determine whether the keyword is in the speech.