Reference 4 stops to get here
Text-to-Speech
Synthesizing natural-sounding speech from text, using neural vocoders and attention-based models.
Your route here
4 stops · basics first
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Feature ✓ understood
A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.
- Feature Engineering ✓ understood
The process of selecting, transforming, and creating input features to improve model performance.
- Audio Processing ✓ understood
Techniques for analyzing, transforming, and understanding audio signals for tasks like speech recognition and music generation.
- Text-to-Speech · you are here ✓ understood
Where it sits
Explore nearby
Language & LLMs Speech Recognition Converting spoken language into text using acoustic models and language models, now dominated by deep learning. Language & LLMs Autoregressive Model A model that generates output one token at a time, using previously generated tokens as input for the next prediction. Language & LLMs Sequence-to-Sequence Models that transform input sequences to output sequences, used for translation, summarization, and generation. Vision & Multimodal Multimodal Model Models processing multiple data types (text, images, audio) jointly, like GPT-4V, Gemini, or CLIP. Language & LLMs Transformer A neural network architecture, introduced in 2017, built from stacked self-attention and feed-forward layers; the basis of nearly every modern large language model.
In the research
All papers →3 papers that build on Text-to-Speech .