Natural language processing (NLP) is the part of AI that works with human language: reading it, sorting it, translating it, answering questions about it and writing it. Spam filters, search engines, voice assistants, machine translation and chatbots are all NLP systems.
What makes it hard is that language is ambiguous and open-ended. “I saw her duck” has two meanings, the same request can be phrased countless ways, and new words keep appearing. Computers work on numbers, so every NLP system has to turn text into numbers, compute on them, and turn the result back into something useful. The pipeline in the diagram is that round trip, and each stage is a design decision: how to cut text into tokens, how to represent each token, and what model sits in the middle.
Three eras
- Rules. Early systems were hand-written grammars and pattern-matching rules. They were exact on the cases their authors anticipated and brittle on everything else.
- Statistics. From the 1990s, systems learned from counts over large text collections: n-gram language models, probabilistic part-of-speech taggers, statistical translation.
- Neural networks. Words became learned vectors called embeddings. In 2013, word2vec showed that good word vectors could be learned from 1.6 billion words in under a day. In 2017 the transformer made it practical to train far larger models on raw text, which led to large language models.
The classic tasks
- Understanding: text classification (spam or not, positive or negative review), picking out the names of people and places, tagging grammar.
- Transforming: machine translation, summarization, rewriting in another style.
- Answering and generating: question answering, dialogue, drafting text and code.
For decades each of these tasks had its own specialised model, trained on its own labelled data. Today a single pre-trained language model, given instructions in plain words, can handle most of them. Much of the work in NLP has moved from building task-specific models to adapting, evaluating and grounding general ones.
The catch
Models learn language from data, so they inherit its gaps. They’re strongest in languages with lots of text online and weaker in the rest, and they pick up the biases of whatever they read. Fluent output is also not proof of understanding: a system can produce a grammatical, confident sentence that is simply false.