Skip to main content
ScienceVerse
Artificial IntelligenceDifficulty 2-4

Language models

A language model is a system that assigns probabilities to sequences of words — in practice, it repeatedly predicts a probability for each possible next token (a word or piece of a word) given the text so far. Large language models are neural networks with billions of parameters trained on vast amounts of text, and they can carry out many tasks described in plain language. Because they generate statistically likely text rather than looking facts up, they can state false things confidently ('hallucination'), and they reproduce biases in their training data.

Sign in to save this concept.

Scaling, adaptation and alignment

Autoregressive LMs are trained by maximum likelihood on next-token prediction. GPT-3 (175B parameters) showed that, at scale, tasks can be specified in the prompt with a few demonstrations and no gradient updates. Kaplan et al. (preprint) reported power-law relationships between loss and model size, data and compute over several orders of magnitude.

  • Instruction tuning and RLHF: supervised fine-tuning on demonstrations, then reinforcement learning from human preference rankings (Ouyang et al., 2022).
  • Confabulation: confidently stated false content; on AA-Omniscience (AI Index 2026), measured rates across 26 models spanned 22%–94%.
  • Bias: models retain biases of their training data (Brown et al., 2020).
  • Contamination: benchmark items in web-scale training corpora can inflate evaluation scores.
Common misconception: Lower cross-entropy loss with scale does not imply factual reliability: scaling laws describe average next-token prediction, while hallucination is measured on specific factual tasks and varies widely between models of similar scale.
Full explanation — the complete reference version every reading depth is based on

Predicting the next word

A goal of statistical language modelling is to learn the probability of sequences of words. Any sequence's probability can be split into a chain of next-word predictions: the probability of the first word, times the probability of the second given the first, and so on. Modern large language models (LLMs) do exactly this, predicting the next token or word in a sentence. To write text, the model picks a next token from its predicted probabilities, adds it to the text, and repeats.

P(w1,…,wn)=∏t=1nP(wt∣w<t)P(w_1, \ldots, w_n) = \prod_{t=1}^{n} P(w_t \mid w_{<t})

The chain rule of probability: a sequence's probability is the product of each word's probability given all the words before it.

Worked example: scoring a continuation

Suppose (with made-up numbers for illustration) a model says that after 'The cat sat on the', the next token is 'mat' with probability 0.6, 'floor' 0.25 and something else 0.15. After 'The cat sat on the mat' it gives a full stop probability 0.5. The probability of continuing with 'mat.' is 0.6 × 0.5 = 0.3. Notice that nothing in this calculation checks whether a cat really sat on a mat — the model scores how likely text is, not whether it is true.

Tokens and word vectors

  • Tokens: models work on pieces of text from a fixed vocabulary. Splitting rare words into subword units (for example with byte pair encoding) lets a model handle words it has never seen whole.
  • Distributed representations: since Bengio and colleagues (2003), neural language models represent each word as a vector of numbers, so similar words get similar vectors and an unseen sentence made of familiar-ish words can still get a sensible probability.
  • Architecture: large models such as GPT-3 are built from transformer layers, the de facto standard architecture for language processing (see Transformers).

From predicting text to following instructions

GPT-3 (2020) had 175 billion parameters and could attempt new tasks from a description and a few examples given as text, without retraining. Scaling studies reported that a language model's loss falls smoothly, as a power law, as model size, data and compute grow. Raw next-word prediction does not by itself make a model helpful: Ouyang and colleagues (2022) fine-tuned GPT-3 with human feedback, and people preferred the answers of a 1.3-billion-parameter tuned model to the 175-billion-parameter original — though it still made simple mistakes.

Limitations you should know

  • Confabulation ('hallucination'): NIST describes the production of confidently stated but false content as a natural result of how generative models are designed, and a survey of the research literature (Ji et al.) finds deep-learning text generation prone to hallucinating unintended text. On one 2026 benchmark of 6,000 factual questions, measured hallucination rates across 26 models ranged from 22% to 94% (AI Index 2026).
  • Bias: GPT-3's authors reported the model retains the biases of its training data and may produce stereotyped or prejudiced content.
  • No built-in fact store: a model's knowledge is whatever patterns its training text contained, up to the date that text was collected.
Common misconception: Misconception: 'A language model looks up answers in a database of facts.' Correct idea: it predicts likely next tokens from patterns in its training text. That is why it can produce fluent, confident, wrong answers — a confident tone is not evidence of accuracy.
Common misconception: Misconception: 'Bigger models don't hallucinate.' Correct idea: scale improves many measures, but hallucination rates on the AI Index 2026's benchmark still varied from 22% to 94% across leading models. Important answers should be checked against reliable sources.
Info: Where this connects: neural networks and transformers explain how the probabilities are computed; training data explains where biases come from; AI agents extend language models with tools and actions.

Ask ScienceVerse

Still curious about Language models? Ask a question, get hints, take a short lesson or try a challenge. The tutor answers only from this concept's approved sources, and says so when it has none.

Ask the tutor about this concept on the full tutor page.

Connections

Prerequisites

Understand these first:

Guided learning path

See everything to learn before this, in order, with your progress:

Leads to

This concept is a building block for:

Related concepts

Check your understanding

Take a quick check of two to five questions, with an explanation for every answer:

See the neighbourhood of Language models in the Knowledge Galaxy

Sources and methodology

  • A goal of statistical language modelling is to learn the joint probability function of sequences of words in a language. (awaiting scientific review)
  • Bengio and colleagues (2003) proposed learning a distributed representation for each word, so that a never-before-seen sequence of words can receive high probability if it is made of words similar to those in sentences already seen. (awaiting scientific review)
  • NIST explains that generative AI models produce outputs that approximate the statistical distribution of their training data; for example, large language models predict the next token or word in a sentence or phrase. (awaiting scientific review)
  • NIST defines 'confabulation' — colloquially called 'hallucination' or 'fabrication' — as the production of confidently stated but erroneous or false content, and describes it as a natural result of the way generative models are designed. (awaiting scientific review)
  • A survey by Ji and colleagues in ACM Computing Surveys reports that deep-learning-based natural language generation is prone to hallucinating unintended text, which degrades system performance and fails to meet user expectations in many real-world scenarios. (awaiting scientific review)
  • Sennrich and colleagues (2016) showed that encoding rare and unknown words as sequences of subword units, including a segmentation based on byte pair encoding, lets neural translation models handle an open vocabulary. (awaiting scientific review)
  • GPT-3, described in 2020, is an autoregressive language model with 175 billion parameters that was evaluated without any gradient updates or fine-tuning, with tasks and a few demonstrations given purely as text. (awaiting scientific review)
  • The GPT-3 authors reported that the model retains the biases of the data it was trained on, which may lead it to generate stereotyped or prejudiced content. (awaiting scientific review)
  • Kaplan and colleagues (2020, preprint) reported that language-model loss falls as a power law in model size, dataset size and training compute, with some trends spanning more than seven orders of magnitude. (awaiting scientific review)
  • Ouyang and colleagues (2022) fine-tuned GPT-3 using human feedback and found human evaluators preferred outputs of a 1.3-billion-parameter InstructGPT model to those of the 175-billion-parameter GPT-3, although InstructGPT still makes simple mistakes. (awaiting scientific review)
  • The Stanford AI Index 2026 reports that on one 6,000-question knowledge and hallucination benchmark (AA-Omniscience), measured hallucination rates across 26 models ranged from 22% to 94%. (awaiting scientific review)
  • Worked calculation (author's own, with illustrative probabilities rather than a real model's): if a model gives 'mat' probability 0.6 after 'The cat sat on the' and then gives '.' probability 0.5 after 'The cat sat on the mat', the probability of the two-token continuation 'mat.' is 0.6 × 0.5 = 0.3. (awaiting scientific review)

Claims marked “awaiting scientific review” cite the sources listed but have not yet been signed off by a scientific reviewer.

Content status: published 1 October 2026.

  • Scientific review: this version has not yet been signed off by a scientific reviewer.
  • The Advanced explanation has not yet been reviewed for age suitability.