Language models
A language model is a system that assigns probabilities to sequences of words — in practice, it repeatedly predicts a probability for each possible next token (a word or piece of a word) given the text so far. Large language models are neural networks with billions of parameters trained on vast amounts of text, and they can carry out many tasks described in plain language. Because they generate statistically likely text rather than looking facts up, they can state false things confidently ('hallucination'), and they reproduce biases in their training data.
Scaling, adaptation and alignment
Autoregressive LMs are trained by maximum likelihood on next-token prediction. GPT-3 (175B parameters) showed that, at scale, tasks can be specified in the prompt with a few demonstrations and no gradient updates. Kaplan et al. (preprint) reported power-law relationships between loss and model size, data and compute over several orders of magnitude.
- Instruction tuning and RLHF: supervised fine-tuning on demonstrations, then reinforcement learning from human preference rankings (Ouyang et al., 2022).
- Confabulation: confidently stated false content; on AA-Omniscience (AI Index 2026), measured rates across 26 models spanned 22%–94%.
- Bias: models retain biases of their training data (Brown et al., 2020).
- Contamination: benchmark items in web-scale training corpora can inflate evaluation scores.
Full explanation — the complete reference version every reading depth is based on
Predicting the next word
A goal of statistical language modelling is to learn the probability of sequences of words. Any sequence's probability can be split into a chain of next-word predictions: the probability of the first word, times the probability of the second given the first, and so on. Modern large language models (LLMs) do exactly this, predicting the next token or word in a sentence. To write text, the model picks a next token from its predicted probabilities, adds it to the text, and repeats.
The chain rule of probability: a sequence's probability is the product of each word's probability given all the words before it.
Worked example: scoring a continuation
Suppose (with made-up numbers for illustration) a model says that after 'The cat sat on the', the next token is 'mat' with probability 0.6, 'floor' 0.25 and something else 0.15. After 'The cat sat on the mat' it gives a full stop probability 0.5. The probability of continuing with 'mat.' is 0.6 × 0.5 = 0.3. Notice that nothing in this calculation checks whether a cat really sat on a mat — the model scores how likely text is, not whether it is true.
Tokens and word vectors
- Tokens: models work on pieces of text from a fixed vocabulary. Splitting rare words into subword units (for example with byte pair encoding) lets a model handle words it has never seen whole.
- Distributed representations: since Bengio and colleagues (2003), neural language models represent each word as a vector of numbers, so similar words get similar vectors and an unseen sentence made of familiar-ish words can still get a sensible probability.
- Architecture: large models such as GPT-3 are built from transformer layers, the de facto standard architecture for language processing (see Transformers).
From predicting text to following instructions
GPT-3 (2020) had 175 billion parameters and could attempt new tasks from a description and a few examples given as text, without retraining. Scaling studies reported that a language model's loss falls smoothly, as a power law, as model size, data and compute grow. Raw next-word prediction does not by itself make a model helpful: Ouyang and colleagues (2022) fine-tuned GPT-3 with human feedback, and people preferred the answers of a 1.3-billion-parameter tuned model to the 175-billion-parameter original — though it still made simple mistakes.
Limitations you should know
- Confabulation ('hallucination'): NIST describes the production of confidently stated but false content as a natural result of how generative models are designed, and a survey of the research literature (Ji et al.) finds deep-learning text generation prone to hallucinating unintended text. On one 2026 benchmark of 6,000 factual questions, measured hallucination rates across 26 models ranged from 22% to 94% (AI Index 2026).
- Bias: GPT-3's authors reported the model retains the biases of its training data and may produce stereotyped or prejudiced content.
- No built-in fact store: a model's knowledge is whatever patterns its training text contained, up to the date that text was collected.
Ask ScienceVerse
Still curious about Language models? Ask a question, get hints, take a short lesson or try a challenge. The tutor answers only from this concept's approved sources, and says so when it has none.
Ask the tutor about this concept on the full tutor page.
Connections
Guided learning path
See everything to learn before this, in order, with your progress:
Related concepts
- Transformers — Application of
- Reinforcement learning — Application of
Check your understanding
Take a quick check of two to five questions, with an explanation for every answer:
See the neighbourhood of Language models in the Knowledge Galaxy
Sources and methodology
- A goal of statistical language modelling is to learn the joint probability function of sequences of words in a language. (awaiting scientific review)
- A Neural Probabilistic Language Model — Peer-reviewed paper
- Bengio and colleagues (2003) proposed learning a distributed representation for each word, so that a never-before-seen sequence of words can receive high probability if it is made of words similar to those in sentences already seen. (awaiting scientific review)
- A Neural Probabilistic Language Model — Peer-reviewed paper
- NIST explains that generative AI models produce outputs that approximate the statistical distribution of their training data; for example, large language models predict the next token or word in a sentence or phrase. (awaiting scientific review)
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (July 2024) — Government or standards body
- NIST defines 'confabulation' — colloquially called 'hallucination' or 'fabrication' — as the production of confidently stated but erroneous or false content, and describes it as a natural result of the way generative models are designed. (awaiting scientific review)
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (July 2024) — Government or standards body
- A survey by Ji and colleagues in ACM Computing Surveys reports that deep-learning-based natural language generation is prone to hallucinating unintended text, which degrades system performance and fails to meet user expectations in many real-world scenarios. (awaiting scientific review)
- Survey of Hallucination in Natural Language Generation — Peer-reviewed paper
- Sennrich and colleagues (2016) showed that encoding rare and unknown words as sequences of subword units, including a segmentation based on byte pair encoding, lets neural translation models handle an open vocabulary. (awaiting scientific review)
- Neural Machine Translation of Rare Words with Subword Units — Peer-reviewed paper
- GPT-3, described in 2020, is an autoregressive language model with 175 billion parameters that was evaluated without any gradient updates or fine-tuning, with tasks and a few demonstrations given purely as text. (awaiting scientific review)
- Language Models are Few-Shot Learners — Peer-reviewed paper
- The GPT-3 authors reported that the model retains the biases of the data it was trained on, which may lead it to generate stereotyped or prejudiced content. (awaiting scientific review)
- Language Models are Few-Shot Learners — Peer-reviewed paper
- Kaplan and colleagues (2020, preprint) reported that language-model loss falls as a power law in model size, dataset size and training compute, with some trends spanning more than seven orders of magnitude. (awaiting scientific review)
- Scaling Laws for Neural Language Models — Other (unclassified)
- Ouyang and colleagues (2022) fine-tuned GPT-3 using human feedback and found human evaluators preferred outputs of a 1.3-billion-parameter InstructGPT model to those of the 175-billion-parameter GPT-3, although InstructGPT still makes simple mistakes. (awaiting scientific review)
- Training language models to follow instructions with human feedback — Peer-reviewed paper
- The Stanford AI Index 2026 reports that on one 6,000-question knowledge and hallucination benchmark (AA-Omniscience), measured hallucination rates across 26 models ranged from 22% to 94%. (awaiting scientific review)
- Artificial Intelligence Index Report 2026 — Other (unclassified)
- Worked calculation (author's own, with illustrative probabilities rather than a real model's): if a model gives 'mat' probability 0.6 after 'The cat sat on the' and then gives '.' probability 0.5 after 'The cat sat on the mat', the probability of the two-token continuation 'mat.' is 0.6 × 0.5 = 0.3. (awaiting scientific review)
- A Neural Probabilistic Language Model — Peer-reviewed paper
Claims marked “awaiting scientific review” cite the sources listed but have not yet been signed off by a scientific reviewer.
Content status: published 1 October 2026.
- Scientific review: this version has not yet been signed off by a scientific reviewer.
- The Advanced explanation has not yet been reviewed for age suitability.