- Number
- 20180053514
- Published
- 2018-02-22
- Filed
- 2016-08-22
- Assignee
- DISNEY ENTERPRISES, INC.
- Inventors
- Singh; Rita et al.
- CPC
- G10L25/15; G10L15/02; G10L17/26; G10L25/51
- Verdict
- Set aside speech-based child age estimation, analytics
- Source
- Google Patents · FreePatentsOnline
Abstract
There is provided a system comprising a microphone, configured to receive an input speech from an individual, an analog-to-digital (A/D) converter to convert the input speech to digital form and generate a digitized speech, a memory storing an executable code and an age estimation database, a hardware processor executing the executable code to receive the digitized speech, identify a plurality of boundaries in the digitized speech delineating a plurality of phonemes in the digitized speech, extract a plurality of formant-based feature vectors from each phoneme in the digitized speech based on at least one of a formant position, a formant bandwidth, and a formant dispersion, compare the plurality of formant-based feature vectors with age determinant formant-based feature vectors of the age estimation database, determine the age of the individual when the comparison finds a match in the age estimation database, and communicate an age-appropriate response to the individual.
Background
BACKGROUND
Advances in voice recognition technology have made voice activated and voice controlled technology more common. Mobile phones and in-home devices now include the ability to listen to speech, respond to activation commands, and execute actions based on voice input. Additionally, an increasing number of voice-controlled and interactive devices may be found in public, such as interacting with guests in theme parks. However, current technology does not enable these voice activated and voice controlled devices to properly estimate age of a speaker based on his or her speech.SUMMARY
The present disclosure is directed to systems and methods for estimating age of a speaker based on speech, substantially as shown in and/or described in connection with at least one of the figures, as set forth more completely in the claims.
Claims
1. A system comprising: a microphone configured to receive an input speech from an individual; an analog-to-digital (A/D) converter configured to convert the input speech from an analog form to a digital form and generate a digitized speech; a memory storing an executable code and an age estimation database including a plurality of age determinant formant-based feature vectors; a hardware processor executing the executable code to: receive the digitized speech from the A/D converter; identify a plurality of boundaries between a plurality of phonemes in the digitized speech; extract a plurality of formant-based feature vectors from one or more phonemes of the plurality of phonemes delineated by the plurality of boundaries, based on at least one of a formant position, a formant bandwidth, and a formant dispersion, wherein the formant dispersion is a geometric mean of the formant spacing; compare the plurality of formant-based feature vectors with the age determinant formant-based feature vectors of the age estimation database; estimate the age of the individual when the comparison finds a match in the age estimation database; and communicate an age-appropriate response to the individual based on the estimated age of the individual.
9. A method for use with a system having a microphone, an analog-to-digital (A/D) converter, a memory storing an executable code, and a hardware processor, the method comprising: receiving, using the hardware processor, a digitized speech from the A/D converter; identifying, using the hardware processor, a plurality of boundaries between a plurality of phonemes in the digitized speech; extracting, using the hardware processor, a plurality of formant-based feature vectors from one or more phonemes of the plurality of phonemes delineated by the plurality of boundaries, based on at least one of a formant position, a formant bandwidth, and a formant dispersion, wherein the formant dispersion is a geometric mean of the formant spacing; comparing, using the hardware processor, the plurality of formant-based feature vectors with the age determinant formant-based feature vectors of the age estimation database; estimating, using the hardware processor, the age of the individual when the comparison finds a match in the age estimation database; and communicating, using the hardware processor, an age-appropriate response to the individual based on the estimated age of the individual.