What technology area does this patent fall under?

Primary CPC classification G10L13/08. Mapped technology areas include Physics.

When was this patent published?

Publication date Thu Nov 02 2023 00:00:00 GMT+0000 (Coordinated Universal Time) (A1). Legal status and post-grant events are not shown on this page.

What related patents are in patentsdb?

We list 8 related publications on this page (citations in our corpus or others sharing the same primary CPC).

Artificial intelligence-based text-to-speech system and method

US2023351999A1 · US · A1

Patent metadata
Field	Value
Publication number	US-2023351999-A1
Application number	US-202318346657-A
Country	US
Kind code	A1
Filing date	Jul 3, 2023
Priority date	May 18, 2017
Publication date	Nov 2, 2023
Grant date	—

How to read this patent

A practical reading order for non-experts. Skip the full description unless you need deep technical detail.

Title
What the patent document calls the invention.
Abstract
A short plain-language summary of the technical disclosure.
Assignees and inventors
Who owns or filed the patent and who is credited as inventor.
Key dates
Filing, priority, publication, and grant dates set the timeline.
First independent claim
The legal scope of protection — read this for what is actually claimed.
CPC / IPC classifications
Technology tags used to group this patent with similar filings.
Citations and related patents
Prior art links and similar publications in this corpus.

Abstract

Official abstract text for this publication.

A technique improves training and speech quality of a text-to-speech (TTS) system having an artificial intelligence, such as a neural network. The TTS system is organized as a front-end subsystem and a back-end subsystem. The front-end subsystem is configured to provide analysis and conversion of text into input vectors, each having at least a base frequency, f 0 , a phenome duration, and a phoneme sequence that is processed by a signal generation unit of the back-end subsystem. The signal generation unit includes the neural network interacting with a pre-existing knowledgebase of phenomes to generate audible speech from the input vectors. The technique applies an error signal from the neural network to correct imperfections of the pre-existing knowledgebase of phenomes to generate audible speech signals. A back-end training system is configured to train the signal generation unit by applying psychoacoustic principles to improve quality of the generated audible speech signal.

First claim

Opening claim text (preview).

What is claimed is: 1 . A text-to-speech (TTS) system including one or more processors and one or more memories configured to perform operations for converting text into a corrected speech signal comprising: training a neural network based upon, at least in part, data of previously generated speech in a pre-existing knowledgebase of phonemes, wherein the previously generated speech has an inaccuracy; generating a lossy representation of at least a portion of the data for use in the training; and applying lossy representation of at least the portion of the data to the previously generated speech for correcting the inaccuracy of the previously generated speech in the pre-existing knowledgebase of phonemes. 2 . The TTS system of claim 1 wherein the lossy representation reduces a representation of a phoneme in the pre-existing knowledgebase of phonemes resulting in a lossy representation of the phoneme where the inaccuracy represents inaudible errors. 3 . The TTS system of claim 1 wherein generating the lossy representation includes limiting derivations of a principal component analysis for at least the portion of the data. 4 . The TTS system of claim 1 wherein the lossy representation is a domain-to-frequency domain transformation of at least the portion of the data. 5 . The TTS system of claim 1 wherein applying the lossy representation includes correcting voiced phonemes of the pre-existing knowledgebase of phonemes using principal component analysis. 6 . The TTS system of claim 1 wherein applying the lossy representation includes correcting unvoiced phonemes of the pre-existing knowledgebase of phonemes using noise band/energy band thresholding. 7 . The TTS system of claim 1 wherein applying the lossy representation includes combining a limited number of frequency bands with specified band widths. 8 . A method of processing text-to-speech (TTS) comprising: training a neural network based upon, at least in part, data of previously generated speech in a pre-existing knowledgebase of phonemes, wherein the previously generated speech has an inaccuracy; generating a lossy representation of at least a portion of the data for use in the training; and applying lossy representation of at least the portion of the data to the previously generated speech for correcting the inaccuracy of the previously generated speech in the pre-existing knowledgebase of phonemes. 9 . The method of claim 8 wherein the lossy representation reduces a representation of a phoneme in the pre-existing knowledgebase of phonemes resulting in a lossy representation of the phoneme where the inaccuracy represents inaudible errors. 10 . The method of claim 8 wherein generating the lossy representation includes limiting derivations of a principal component analysis for at least the portion of the data. 11 . The method of claim 8 wherein the lossy representation is a domain-to-frequency domain transformation of at least the portion of the data. 12 . The method of claim 8 wherein applying the lossy representation includes correcting voiced phonemes of the pre-existing knowledgebase of phonemes using principal component analysis. 13 . The method of claim 8 wherein applying the lossy representation includes correcting unvoiced phonemes of the pre-existing knowledgebase of phonemes using noise band/energy band thresholding. 14 . The method of claim 8 wherein applying the lossy representation includes combining a limited number of frequency bands with specified band widths. 15 . A non-transitory computer-readable medium having program instructions which, when executed across one or more processors, causes at least a portion of the one or more processors to perform operations comprising: training a neural network based upon, at least in part, data of previously generated speech in a pre-existing knowledgebase of phonemes, wherein the previously generated speech has an inaccuracy; generating a lossy representation of at least a portion of the data for use in the training; and applying lossy representation of at least the portion of the data to the previously generated speech for correcting the inaccuracy of the previously generated speech in the pre-existing knowledgebase of phonemes. 16 . The non-transitory computer-readable medium of claim 15 wherein the lossy representation reduces a representation of a phoneme in the pre-existing knowledgebase of phonemes resulting in a lossy representation of the phoneme where the inaccuracy represents inaudible errors. 17 . The non-transitory computer-readable medium of claim 15 wherein generating the lossy representation includes limiting derivations of a principal component analysis for at least the portion of the data. 18 . The non-transitory computer-readable medium of claim 15 wherein applying the lossy representation includes correcting voiced phonemes of the pre-existing knowledgebase of phonemes using principal component analysis. 19 . The non-transitory computer-readable medium of claim 15 wherein applying the lossy representation includes correcting unvoiced phonemes of the pre-existing knowledgebase of phonemes using noise band/energy band thresholding. 20 . The non-transitory computer-readable medium of claim 15 wherein applying the lossy representation includes combining a limited number of frequency bands with specified band widths.

Assignees

Telepathy Labs Inc

Inventors

Classifications

G06N3/08
Learning methods · CPC title
G10L25/30
using neural networks · CPC title
G06N3/0499
Feedforward networks · CPC title
G06N3/09
Supervised learning · CPC title
G06N3/0495
Quantised networks; Sparse networks; Compressed networks · CPC title

Patent family

Related publications grouped by family.

View patent family 64271935

External sources

Frequently asked questions

Answers are generated from the same data shown on this page.

What does patent US2023351999A1 cover?: A technique improves training and speech quality of a text-to-speech (TTS) system having an artificial intelligence, such as a neural network. The TTS system is organized as a front-end subsystem and a back-end subsystem. The front-end subsystem is configured to provide analysis and conversion of text into input vectors, each having at least a base frequency, f 0 , a phenome duration, and a pho…
Who is the assignee on this patent?: Telepathy Labs Inc
What technology area does this patent fall under?: Primary CPC classification G10L13/08. Mapped technology areas include Physics.
When was this patent published?: Publication date Thu Nov 02 2023 00:00:00 GMT+0000 (Coordinated Universal Time) (A1). Legal status and post-grant events are not shown on this page.
What related patents are in patentsdb?: We list 8 related publications on this page (citations in our corpus or others sharing the same primary CPC).