Standard 3 stops to get here
Mixture of Experts
An architecture where multiple specialized sub-networks (experts) process inputs, with a gating network routing to relevant experts.
Your route here
3 stops · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Neural Network ✓ understood
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
- Feedforward Network ✓ understood
A neural network where information flows in one direction from input to output without cycles.
- Mixture of Experts · you are here ✓ understood
Where it sits
Mixture of Experts
Leads to
Nothing yet: a destination in its own right.
Explore nearby
Foundations Ensemble Learning Combining multiple models to produce better predictions than any individual model (bagging, boosting, stacking). Language & LLMs Transformer A neural network architecture, introduced in 2017, built from stacked self-attention and feed-forward layers; the basis of nearly every modern large language model. Language & LLMs Large Language Model A neural network, almost always a transformer, trained on vast amounts of text to predict the next token, which lets it write, answer, summarize and follow instructions. Language & LLMs Neural Scaling Laws Empirical relationships showing how model performance improves predictably with model size, data, and compute. Training Data Parallelism Replicating the model across devices, each processing different data batches.
In the research
All papers →6 papers that build on Mixture of Experts ; showing 5, canon first.