What technology area does this patent fall under?

Primary CPC classification G10L13/08. Mapped technology areas include Physics.

When was this patent published?

Publication date Tue Aug 06 2019 00:00:00 GMT+0000 (Coordinated Universal Time) (B2). Legal status and post-grant events are not shown on this page.

What related patents are in patentsdb?

We list 3 related publications on this page (citations in our corpus or others sharing the same primary CPC).

Artificial intelligence-based text-to-speech system and method

US10373605B2 · US · B2

Patent metadata
Field	Value
Publication number	US-10373605-B2
Application number	US-201816022823-A
Country	US
Kind code	B2
Filing date	Jun 29, 2018
Priority date	May 18, 2017
Publication date	Aug 6, 2019
Grant date	Aug 6, 2019

How to read this patent

A practical reading order for non-experts. Skip the full description unless you need deep technical detail.

Title
What the patent document calls the invention.
Abstract
A short plain-language summary of the technical disclosure.
Assignees and inventors
Who owns or filed the patent and who is credited as inventor.
Key dates
Filing, priority, publication, and grant dates set the timeline.
First independent claim
The legal scope of protection — read this for what is actually claimed.
CPC / IPC classifications
Technology tags used to group this patent with similar filings.
Citations and related patents
Prior art links and similar publications in this corpus.

Abstract

Official abstract text for this publication.

A technique improves training and speech quality of a text-to-speech (TTS) system having an artificial intelligence, such as a neural network. The TTS system is organized as a front-end subsystem and a back-end subsystem. The front-end subsystem is configured to provide analysis and conversion of text into input vectors, each having at least a base frequency, f0, a phenome duration, and a phoneme sequence that is processed by a signal generation unit of the back-end subsystem. The signal generation unit includes the neural network interacting with a pre-existing knowledgebase of phenomes to generate audible speech from the input vectors. The technique applies an error signal from the neural network to correct imperfections of the pre-existing knowledgebase of phenomes to generate audible speech signals. A back-end training system is configured to train the signal generation unit by applying psychoacoustic principles to improve quality of the generated audible speech signals.

First claim

Opening claim text (preview).

What is claimed is: 1. A text-to-speech (TTS) training system comprising: a subsystem configured to receive an input vector from conversion of text, the subsystem including a neural network interacting with a pre-existing knowledgebase of phonemes to apply an error signal to correct for speech signal distortions of the pre-existing knowledgebase of phonemes to generate a corrected speech signal; and a training subsystem coupled to the subsystem, the training subsystem configured to iteratively correct the subsystem for the speech signal distortions of the pre-existing knowledgebase of phonemes based on psychoacoustic processing for the subsystem to apply the error signal to correct for the speech signal distortions of the pre-existing knowledgebase of phonemes to generate the corrected speech signal, wherein the training subsystem is further configured to ignore inaudible errors of the corrected speech signal based on masking. 2. The TTS training system of claim 1 wherein the training subsystem is further configured to calculate audible errors of the corrected speech signal by the psychoacoustic processing. 3. The TTS training system of claim 1 wherein masking includes one of frequency masking and temporal masking. 4. The TTS training system of claim 1 wherein the training subsystem is further configured to iteratively modify the neural network to correct the subsystem. 5. The TTS training system of claim 1 further comprising a psychoacoustic generator configured to analyze a reference audio signal to determine masking information. 6. The TTS training system of claim 5 wherein the psychoacoustic generator is further configured to identify locations and energy levels that are audible and inaudible. 7. The TTS training system of claim 1 further comprising a quality indicator calculator configured to determine a quality indicator based on audible errors using the psychoacoustic processing to ignore inaudible errors. 8. The TTS training system of claim 7 wherein the quality indicator is calculated based on a total of audible error signal energy. 9. The TTS training system of claim 7 wherein the iterative correction further comprises a modification of the neural network that is performed so that a total audible error signal energy is below a quality threshold. 10. The TTS training system of claim 9 wherein the quality indicator at least converges close to zero with the modification of the neural network. 11. A method of training text-to-speech (TTS) processing comprising: receiving, by a subsystem, an input vector from conversion of text; interacting, by a neural network of the subsystem, with a pre-existing knowledgebase of phonemes to apply an error signal to correct for speech signal distortions of the pre-existing knowledgebase of phonemes to generate a corrected speech signal; iteratively correcting, by a training subsystem coupled to the subsystem, the subsystem for the speech signal distortions of the pre-existing knowledgebase of phonemes based on psychoacoustic processing for the subsystem to apply the error signal to correct for the speech signal distortions of the pre-existing knowledgebase of phonemes to generate the corrected speech signal; and ignoring, by the training subsystem, inaudible errors of the corrected speech signal based on masking. 12. The method of training TTS processing of claim 11 further comprising calculating audible errors of the corrected speech signal by the psychoacoustic processing. 13. The method of training TTS processing of claim 11 wherein masking includes one of frequency masking and temporal masking. 14. The method of training TTS processing of claim 11 wherein the iterative correction further comprises modifying the neural network to correct the subsystem. 15. The method of training TTS processing of claim 11 further comprising analyzing a reference audio signal to determine masking information. 16. The method of training TTS processing of claim 11 further comprising identifying locations and energy levels that are audible and inaudible. 17. The method of training TTS processing of claim 11 further comprising: determining a quality indicator based on audible errors; and ignoring inaudible errors using the psychoacoustic processing. 18. The method of training TTS processing of claim 17 wherein the quality indicator is calculated based on a total of audible error signal energy. 19. The method of training TTS processing of claim 17 wherein the iterative correction further comprises modifying the neural network so that a total audible error signal energy is below a quality threshold. 20. A non-transitory computer-readable medium having program instructions for training text-to-speech (TTS) processing which, when executed across one or more processors, causes at least a portion of the one or more processors to perform operations comprising: receiving, by a subsystem, an input vector from conversion of text; interacting, by a neural network of the subsystem, with a pre-existing knowledgebase of phonemes to apply an error signal to correct for speech signal distortions of the pre-existing knowledgebase of phonemes to generate a corrected speech signal; iteratively correcting, by a training subsystem coupled to the subsystem, the subsystem for the speech signal distortions of the pre-existing knowledgebase of phonemes based on psychoacoustic processing to apply the error signal to correct for the speech signal distortions of the pre-existing knowledgebase of phonemes to generate the corrected speech signal; and ignoring, by the training subsystem, inaudible errors of the corrected speech signal based on masking.

Assignees

Telepathy Labs Inc

Inventors

Classifications

G06F18/2135
based on approximation criteria, e.g. principal component analysis · CPC title
G06F18/217
Validation; Performance evaluation; Active pattern learning techniques · CPC title
G06N3/042
Knowledge-based neural networks; Logical representations of neural networks · CPC title
G10L13/08Primary
Text analysis or generation of parameters for speech synthesis out of text, e.g. grapheme to phoneme translation, prosody generation or stress or intonation determination · CPC title
G10L13/04
Details of speech synthesis systems, e.g. synthesiser structure or memory management · CPC title

Patent family

Related publications grouped by family.

View patent family 64271935

External sources

Frequently asked questions

Answers are generated from the same data shown on this page.

What does patent US10373605B2 cover?: A technique improves training and speech quality of a text-to-speech (TTS) system having an artificial intelligence, such as a neural network. The TTS system is organized as a front-end subsystem and a back-end subsystem. The front-end subsystem is configured to provide analysis and conversion of text into input vectors, each having at least a base frequency, f0, a phenome duration, and a phone…
Who is the assignee on this patent?: Telepathy Labs Inc
What technology area does this patent fall under?: Primary CPC classification G10L13/08. Mapped technology areas include Physics.
When was this patent published?: Publication date Tue Aug 06 2019 00:00:00 GMT+0000 (Coordinated Universal Time) (B2). Legal status and post-grant events are not shown on this page.
What related patents are in patentsdb?: We list 3 related publications on this page (citations in our corpus or others sharing the same primary CPC).

How to read this patent

Abstract

First claim

Assignees

Inventors

Classifications

Patent family

External sources

Related patents

Method for operating a binaural hearing aid system and a binaural hearing aid system

Perceptual power reduction system and method

System and method for automatic detection of abnormal stress patterns in unit selection synthesis

Frequently asked questions