Method and system for acoustic data selection for training the parameters of an acoustic model

US9972306B2 · US · B2

Patent metadata
FieldValue
Publication numberUS-9972306-B2
Application numberUS-201313959171-A
CountryUS
Kind codeB2
Filing dateAug 5, 2013
Priority dateAug 7, 2012
Publication dateMay 15, 2018
Grant dateMay 15, 2018

How to read this patent

A practical reading order for non-experts. Skip the full description unless you need deep technical detail.

  1. Title

    What the patent document calls the invention.

  2. Abstract

    A short plain-language summary of the technical disclosure.

  3. Assignees and inventors

    Who owns or filed the patent and who is credited as inventor.

  4. Key dates

    Filing, priority, publication, and grant dates set the timeline.

  5. First independent claim

    The legal scope of protection — read this for what is actually claimed.

  6. CPC / IPC classifications

    Technology tags used to group this patent with similar filings.

  7. Citations and related patents

    Prior art links and similar publications in this corpus.

Abstract

Official abstract text for this publication.

A system and method are presented for acoustic data selection of a particular quality for training the parameters of an acoustic model, such as a Hidden Markov Model and Gaussian Mixture Model, for example, in automatic speech recognition systems in the speech analytics field. A raw acoustic model may be trained using a given speech corpus and maximum likelihood criteria. A series of operations are performed, such as a forced Viterbi-alignment, calculations of likelihood scores, and phoneme recognition, for example, to form a subset corpus of training data. During the process, audio files of a quality that does not meet a criterion, such as poor quality audio files, may be automatically rejected from the corpus. The subset may then be used to train a new acoustic model.

First claim

Opening claim text (preview).

The invention claimed is: 1. A computer-implemented method for training acoustic models in an automatic speech recognition system through the selection of acoustic data comprising the steps of: a. training a first acoustic model in the automatic speech recognition system using a training-data corpus comprising a plurality of speech audio files and a respective plurality of transcriptions for the plurality of speech audio files; b. performing a forced Viterbi alignment of the plurality of speech audio files using the trained first acoustic model in the automatic speech recognition system and determining an average frame likelihood score β r for each of the plurality of speech audio files; c. calculating a global frame likelihood score δ for the plurality of speech audio files, wherein the global frame likelihood score δ comprises an average of frame likelihoods over the entire corpus; d. performing a phoneme recognition of the plurality of speech audio files using the trained first acoustic model and the plurality of transcriptions in the automatic speech recognition system; e. calculating a phoneme recognition accuracy γ for each of the plurality of speech audio files and a global phoneme recognition accuracy v for the plurality of speech audio files; f. creating a subset training-data corpus comprising audio files retained from the plurality of speech audio files which meet at least one predetermined criterion indicating that an audio file has good audio quality, the at least one predetermined criterion comprising at least one criterion selected from the group comprising: a first criterion based on the average frame likelihood score β of the retained speech audio file and the global frame likelihood score δ; and a second criterion based on the phoneme recognition accuracy γ of the retained speech audio file and the global phoneme recognition accuracy v; and g. training a second acoustic model in the automatic speech recognition system using the subset training-data corpus. 2. The method of claim 1 , wherein step (a) further comprises the steps of: a.1. calculating a maximum likelihood criterion of the training-data corpus; and a.2. estimating parameters of a probability distribution of said first acoustic model that maximize the maximum likelihood criterion. 3. The method of claim 1 , wherein said model comprises a Hidden Markov Model and a Gaussian Mixture Model. 4. The method of claim 1 , wherein step (b) further comprises: obtaining a total likelihood score α r for each of the plurality of speech audio files. 5. The method of claim 4 , wherein α r = p ⁡ ( x 1 ❘ q 1 ) ⁢ ∏ i = 2 N ⁢ ⁢ P ⁡ ( q i ❘ q i - 1 ) ⁢ p ⁡ ( x i ❘ q i ) , where P(q i |q i-1 ) represents a Hidden Markov Model state transition probability between states ‘i−1’ and ‘i’ and p(x i |q i ) represents a state emission likelihood of a feature vector x i being present in a state q i . 6. The method of claim 4 , further comprising using the mathematical equation β r = α r f r to determine the average frame likelihood score of an audio file, wherein β r is the average frame likelihood score, α r is a total likelihood score of the audio file, and f r is a number of feature frames of the audio file. 7. The method of claim 1 , wherein the first criterion comprises determining whether the average frame likelihood β r of the retained audio file satisfies the criterion β r ≧δ+Δ, where Δ is a first predetermined threshold, and wherein the second criterion comprises determining whether the phoneme recognition accuracy γ g of the retained audio file satisfies the criterion γ g ≧v+μ, where μ is a second predetermined threshold. 8. The method of claim 7 , wherein Δ=−0.1δ. 9. The method of claim 7 , wherein μ=−0.2 v. 10. The method of claim 1 further comprising the step of using the mathematical equation δ = ∑ r = 1 R ⁢ ⁢ β r R to obtain the global frame likelihood score δ, wherein β r is the average frame likelihood score and R is the total number of the plurality of speech audio files. 11. The method of claim 1 further comprising the step of using the mathematical equation v = ∑ r = 1 R

Assignees

Inventors

Classifications

Patent family

Related publications grouped by family.

External sources

Frequently asked questions

Answers are generated from the same data shown on this page.

What does patent US9972306B2 cover?
A system and method are presented for acoustic data selection of a particular quality for training the parameters of an acoustic model, such as a Hidden Markov Model and Gaussian Mixture Model, for example, in automatic speech recognition systems in the speech analytics field. A raw acoustic model may be trained using a given speech corpus and maximum likelihood criteria. A series of operations…
Who is the assignee on this patent?
Interactive Intelligence Inc, Interactive Intelligence Group Inc
What technology area does this patent fall under?
Primary CPC classification G10L15/063. Mapped technology areas include Physics.
When was this patent published?
Publication date Tue May 15 2018 00:00:00 GMT+0000 (Coordinated Universal Time) (B2). Legal status and post-grant events are not shown on this page.
What related patents are in patentsdb?
We list 8 related publications on this page (citations in our corpus or others sharing the same primary CPC).