AI agents
An AI agent is a system that perceives its environment, decides what to do and takes actions towards a goal, repeating that loop rather than giving a single answer. Today's best-known agents wrap a language model in a loop with memory, planning and tools — search, calculators, code, a computer's mouse and keyboard — so it can carry out multi-step tasks. Agents have improved quickly on benchmarks (as of the AI Index 2026) but still fail often, and acting in the world raises new risks such as compounding errors and prompt injection, so their autonomy needs limits and human oversight.
Design patterns and evidence
The agent–environment interface from reinforcement learning frames agents generally; language-model agents implement the policy with a pretrained model prompted or fine-tuned to emit actions. ReAct (prompted, one or two examples) beat imitation- and RL-trained baselines by 34 and 10 points of absolute success on ALFWorld and WebShop; Toolformer learned self-supervised API calling for a calculator, QA system, search, translation and calendar.
- Capability (AI Index 2026): best OSWorld accuracy roughly 12% → 66.3%, about 6 points below the human baseline; roughly one in three failures on structured benchmarks.
- Security: indirect prompt injection exploits the blurred boundary between data and instructions (Greshake et al., 2023); NIST notes demonstrated data theft and remote code execution.
- Autonomy: NIST frames AI systems as operating with varying levels of autonomy, so oversight must scale with what an agent can do unaided.
Full explanation — the complete reference version every reading depth is based on
From answers to actions
In reinforcement learning, the agent is the learner and decision-maker, and the environment is everything it interacts with: the agent acts, the environment responds with a new situation, and the loop continues. The same picture describes modern AI agents. What makes something an agent is not intelligence but the loop — observe, decide, act, observe the result — and the fact that its actions change something. NIST notes AI systems operate with varying levels of autonomy; agents sit towards the more autonomous end.
Inside a language-model agent
- Profile: what role the agent is playing and what it is for.
- Memory: what it has seen and done so far in the task (and sometimes across tasks).
- Planning: breaking a goal into steps and revising the plan when something fails.
- Action: turning decisions into concrete outputs — a search query, a calculator call, a click, a line of code.
This four-part breakdown comes from a 2024 survey of language-model-based agents. Two influential designs show the pieces in action. ReAct (2022) has the model alternate between writing out its reasoning and taking actions such as looking something up, then reading the result before continuing. Toolformer (2023) trained a model to decide for itself when to call tools such as a calculator, a search engine, a translator or a calendar, and how to use what comes back.
Worked example: an agent loop
- Goal: 'What is 17% of 2,340?'
- Reason: 'Arithmetic is easy to get wrong by predicting text; I should use the calculator tool.'
- Act: call calculator(0.17 × 2340).
- Observe: the tool returns 397.8.
- Answer: '17% of 2,340 is 397.8.' The answer came from the tool, not from the model guessing digits.
Long tasks are harder. If each step independently succeeds 90% of the time, a five-step task succeeds only 0.9⁵ ≈ 59% of the time. Small error rates compound, which is why checking intermediate results matters.
If steps succeed independently, the chance that all n succeed is the product of their success probabilities.
When every step has the same success probability p, the product becomes p to the power n.
How capable are agents? (as of 2026)
The Stanford AI Index 2026 reported that on OSWorld — real computer tasks across operating systems — the best agent accuracy had risen from roughly 12% to 66.3%, within about 6 percentage points of the human baseline of about 72%. The same report notes agents still failed roughly one in three attempts on structured benchmarks. Earlier, ReAct had outperformed imitation- and reinforcement-learning baselines on two interactive benchmarks by 34 and 10 percentage points of success rate.
Risks of acting in the world
- Prompt injection: because agents read web pages, emails and documents, an attacker can hide instructions in that data. Greshake and colleagues (2023) showed such indirect prompt injection works against real applications, and NIST notes demonstrations that stole data or ran malicious code.
- Compounding errors: one wrong step early in a task can derail everything after it.
- Wrong objectives: like reinforcement-learning agents, a system pursuing a badly specified goal can take actions nobody intended.
Ask ScienceVerse
Still curious about AI agents? Ask a question, get hints, take a short lesson or try a challenge. The tutor answers only from this concept's approved sources, and says so when it has none.
Ask the tutor about this concept on the full tutor page.
Connections
Guided learning path
See everything to learn before this, in order, with your progress:
Related concepts
- Reinforcement learning — Related to
- Robotics — Related to
Check your understanding
Take a quick check of two to five questions, with an explanation for every answer:
Sources and methodology
- In Sutton and Barto's formulation, the learner and decision-maker is called the agent and everything outside it is the environment; the two interact continually, the agent selecting actions and the environment responding by presenting new situations. (awaiting scientific review)
- NIST's AI Risk Management Framework states that AI systems are designed to operate with varying levels of autonomy. (awaiting scientific review)
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (January 2023) — Government or standards body
- A 2024 survey of language-model-based autonomous agents (Wang and colleagues) proposes a unified architecture composed of a profiling module, a memory module, a planning module and an action module, where the action module translates the agent's decisions into specific outputs. (awaiting scientific review)
- A survey on large language model based autonomous agents — Peer-reviewed paper
- ReAct (Yao and colleagues, 2022) has a language model generate reasoning traces and task-specific actions in an interleaved manner, with the actions letting it gather information from external sources such as knowledge bases or environments. (awaiting scientific review)
- ReAct: Synergizing Reasoning and Acting in Language Models — Peer-reviewed paper
- On the ALFWorld and WebShop interactive decision-making benchmarks, ReAct outperformed imitation-learning and reinforcement-learning methods by an absolute success rate of 34 and 10 percentage points respectively, while prompted with only one or two in-context examples. (awaiting scientific review)
- ReAct: Synergizing Reasoning and Acting in Language Models — Peer-reviewed paper
- Toolformer (Schick and colleagues, 2023) is a language model trained to decide which APIs to call, when to call them, what arguments to pass and how to incorporate the results, using tools including a calculator, a question-answering system, two search engines, a translation system and a calendar. (awaiting scientific review)
- Toolformer: Language Models Can Teach Themselves to Use Tools — Peer-reviewed paper
- The Stanford AI Index 2026 reports that on OSWorld, which tests agents on real computer tasks across operating systems, the best agent accuracy rose from roughly 12% to 66.3%, within about 6 percentage points of the human baseline (about 72%), while agents still failed roughly one in three attempts on structured benchmarks. (awaiting scientific review)
- Artificial Intelligence Index Report 2026 — Other (unclassified)
- Greshake and colleagues (2023) showed that adversaries can remotely exploit applications built on language models by planting instructions in data the application is likely to retrieve — 'indirect prompt injection' — because such applications blur the line between data and instructions. (awaiting scientific review)
- NIST's Generative AI Profile notes that security researchers have demonstrated indirect prompt injections that exploit vulnerabilities by stealing proprietary data or running malicious code remotely on a machine. (awaiting scientific review)
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (July 2024) — Government or standards body
- Worked calculation (author's own, assuming each step succeeds independently): if every step of a task succeeds with probability 0.9, a five-step task succeeds with probability 0.9⁵ ≈ 0.59, so small per-step error rates compound over long tasks. (awaiting scientific review)
- Artificial Intelligence Index Report 2026 — Other (unclassified)
Claims marked “awaiting scientific review” cite the sources listed but have not yet been signed off by a scientific reviewer.
Content status: published 1 October 2026.
- Scientific review: this version has not yet been signed off by a scientific reviewer.
- The Advanced explanation has not yet been reviewed for age suitability.