Outer Rim Archives
Archives · 2022 · 11232786

Granted patent

System and method to improve performance of a speech recognition system by measuring amount of confusion between words

Number
11232786
Published
2022-01-25
Filed
2019-11-27
Assignee
DISNEY ENTERPRISES, INC.
Inventors
Tiwari; Sanchita, Shu; Chang
CPC
G10L15/197; G10L15/02; G10L15/183; G10L15/187
Verdict
Set aside generic speech-recognition research
Source
Google Patents · FreePatentsOnline

Abstract

Systems and methods to improve the performance of an automatic speech recognition (ASR) system using a confusion index indicative of the amount of confusion between words are described, where a confusion index (CI) or score is calculated by receiving a first word (Word1) and a second word (Word2), calculating an acoustic score (A12) indicative of the phonetic difference between Word1 and Word2, calculating a weighted language score (W(U1+U2), indicative of a weighted likelihood (or word frequency) of Word1 and Word2 occurring in the corpus, the confusion index CI incorporating both the acoustic score and the weighted language score, such that the CI for words that sound alike and have a high likelihood of occurring in the corpus will be higher than the CI for words that sound alike and do not have a high likelihood of occurring in the corpus. In some embodiments, the CI may be used to artificially boost uncommon words in a corpus to improve their visibility, to add context to uncommon words in a corpus to avoid conflict with common words, and to remove unimportant words from the lexicon to avoid conflicts with other corpus words.

Background

BACKGROUND (1) Significant efforts have been made by companies and academia to measure (or quantify) the confusion between similar sounding words to effectively resolve or minimize the conflicts between them. To date, such measurements have been based on an approach using the acoustic (sound) similarities between the words. However, there are significant challenges with such an acoustic-based approach. (2) While current approaches using acoustic similarities to identify confusing words, the acoustic similarity approach is a “one size fits all” approach, i.e., any two words that have the same acoustic sound differences (sound the same or have the same pronunciation) get treated alike (or get assigned the same level confusion). Such an approach does not provide any insight into the degree or level of confusion between the similar sounding words and how significantly such confusion can affect the models (e.g., language models and the like). (3) Accordingly, it would be desirable to have a system and method that overcomes the shortcomings of the prior art and provides an approach that measures the degree of confusion between words to enable more intelligent handling of similar sounding words to reduce the negative effects on language models and Automatic Speech Recognition (ASR) systems that use the models.

Claims

1. A method for improving the performance of an automatic speech recognition (ASR) system that uses a language model based on a corpus data set which includes words from a generic corpus data set and a domain-specific data set, using a confusion index indicative of the amount of confusion between words from the data sets, comprising: determining the confusion index using a method comprising: receiving a first word from the generic corpus data set and a second word from the domain specific data set; calculating an acoustic score A12 indicative of an acoustic distance between the first word and the second word using a lexicon having a phonetic breakdown of the first word and the second word; calculating a weighted language score indicative of the likelihood of the first word and the second word occurring in the corpus data set, comprising performing an equation: W(U1+U2), where U1 and U2 are unigram values of the first word and the second word, respectively, and W is a weighting factor in the weighted language score; calculating the confusion index (CI) using the acoustic score and the weighted language score, comprising performing an equation: CI=1/(e.sup.A12−e.sup.−W(U1+U1)); and adjusting the corpus data set based on the value of CI. || 3. A method for improving the performance of an automatic speech recognition (ASR) system that uses a language model based on a corpus data set which includes words from a generic corpus data set and words from a domain-specific data set, using a confusion index indicative of the amount of confusion between words from the generic corpus data set and the domain-specific data set, comprising: determining the confusion index using a method comprising: receiving a first word from the domain-specific data set and a second word from the generic corpus data set; calculating an acoustic score indicative of the acoustic distance between the first word and the second word using a lexicon having a phonetic breakdown of the first word and the second word; calculating a weighted language score indicative of the likelihood of the first word and the second word occurring in the corpus data set; wherein the weighted language score comprises an equation: W(U1+U2), where U1 and U2 are unigram values of the first word and the second word, respectively, and W is a weighting; and calculating the confusion index (CI) using the acoustic score and the weighted language score. || 17. A method for improving the performance of an automatic speech recognition (ASR) system that uses a language model based on a corpus data set which includes words from a corpus first data set and a domain-specific second data set, using an amount of confusion between words, comprising: calculating a confusion index (CI), comprising: receiving a first word from the first data set and a second word from the second data set; calculating an acoustic score indicative of the acoustic distance between the first word and the second word using a lexicon having a phonetic breakdown of the first word and the second word; calculating a weighted language score indicative of the likelihood of the first word and the second word occurring in the corpus data set; wherein the weighted language score comprises an equation: W(U1+U2), where U1 and U2 are unigram values of the first word and the second word, respectively, and W is a weighting factor; and calculating the confusion index (CI) using the acoustic score and the weighted language score.