AI agent glossary
The vocabulary around AI agents, explained in plain language.
- Agent loop
The cycle at the centre of every agent: send the task and what has happened so far to the model, run the tool calls it asks for, add the results, and repeat. The loop ends when the model answers without requesting another tool.
- Agentic AI
The umbrella term for AI systems that act instead of only answering: single agents, groups of agents and workflows with agent steps. "AI agent" names one such system; "agentic AI" names the approach.
- AI agent
Software that uses a language model to work toward a goal by choosing and carrying out actions in a loop: it calls a tool, looks at the result and decides the next step, until the task is done or it needs a person. See What is an AI agent?
- Benchmark
A fixed set of tasks with an automatic way to check the result, used to compare agents or models. Scores are only comparable within one version of one benchmark. See benchmarks.
- Bring your own key (BYOK)
An arrangement in which an agent uses an API key you supply for a model provider, so you pay that provider per token, instead of paying the agent's vendor a subscription that includes model usage.
- Browser agent
An agent that completes tasks by controlling a web browser: navigating, reading pages, filling in forms. See browser agents.
- Cloud (asynchronous) agent
An agent that runs a task in a remote environment while you do something else and returns the result later, for a coding agent usually as a pull request.
- Coding agent
An agent that works on software: it reads a codebase, edits files, runs commands and tests, and can open pull requests. See the coding agent rankings.
- Computer use
An agent operating a graphical interface the way a person does: looking at screenshots, moving the pointer, clicking and typing. It works with software that has no API and is slower and less reliable than calling one.
- Context engineering
Deciding what goes into the model's context at each step: which instructions, files, search results and memories to include, and what to leave out or summarise. For agents it usually matters more than the wording of any single prompt.
- Context window
The amount of text a model can take into account in one request, measured in tokens. An agent's instructions, the conversation, file contents and tool results all have to fit, which is why long sessions get slower, costlier and less accurate.
- Evals
Tests for AI behaviour: a set of example tasks and a way to grade the outputs, run again whenever the prompt, the tools or the model changes. An agent team's equivalent of a test suite.
- Guardrails
Checks placed around an agent to keep it within bounds: validating inputs and outputs, limiting which tools and data it can reach, and requiring approval for sensitive actions.
- Hallucination
A confident statement by a model that is not true or not supported by its sources. Agents reduce the risk by checking their work with tools, such as running the code or citing the document.
- Harness
The software around the model that makes it an agent: the loop, the tools, permissions, context management and the interface. Two products can run the same model and behave very differently because their harnesses differ. Also called a scaffold.
- Human in the loop
A design in which the agent pauses for a person to approve, correct or take over at defined points, typically before actions that are costly or hard to undo.
- Memory
Information an agent keeps between sessions, such as preferences, earlier decisions and facts about a project, stored outside the model and loaded back into its context when relevant.
- Merge rate
The share of pull requests opened by a coding agent that end up merged. A rough public signal of how usable its output is, which also depends on how the agent is used. See GitHub activity.
- Model Context Protocol (MCP)
An open standard for connecting agents to tools and data. A service is wrapped once as an MCP server and can then be used by any agent that supports the protocol, instead of needing a separate integration for each agent.
- Multi-agent system
Several agents with different roles or instructions working on one task and passing results to each other, coordinated by an orchestrator or by fixed hand-off rules.
- Orchestration
Coordinating the steps of an agent or several agents: what runs in which order, what state is kept, what happens on failure and where a person is asked to approve.
- Project instruction file (AGENTS.md)
A Markdown file in a repository that tells coding agents how the project works: how to build and test it, conventions and things to avoid.
AGENTS.mdis the common name across tools; some agents use their own, such asCLAUDE.md.- Prompt injection
An attack in which text the agent reads, such as a web page, an email or a document, contains instructions meant to hijack it. It is the main security risk for agents that handle untrusted content and can also take actions.
- Reasoning effort
A setting on many current models that controls how much the model thinks before it answers. Higher effort improves results on hard tasks and costs more time and tokens.
- Resolution rate
In customer support, the share of conversations an AI agent closes without a person. Vendors define "resolved" differently, so their published rates are not directly comparable.
- Retrieval-augmented generation (RAG)
Fetching relevant documents at question time and giving them to the model, so that answers rest on your own content instead of the model's memory. The basis of most support and documentation assistants.
- Sandbox
An isolated environment, such as a container or virtual machine, in which an agent runs commands so that a mistake or a malicious instruction cannot reach the rest of the system.
- Skill
A packaged set of instructions, and sometimes scripts, that an agent loads only when a task calls for it, such as how to produce a certain kind of document. Skills extend an agent without filling its context all the time.
- Sub-agent
A separate agent instance that a main agent starts for a contained piece of work, such as searching a codebase. It has its own context and reports back a summary, which keeps the main agent's context small.
- System prompt
Standing instructions given to the model before the conversation starts: its role, the tools it has, the rules it must follow. In coding agents, a project instruction file is added to it.
- Token
The unit in which models read and write text, roughly three-quarters of an English word. Context limits and API prices are both counted in tokens.
- Tool use (function calling)
The mechanism that lets a model act. The application describes the available tools, the model replies with a structured request to call one, the application runs it and returns the result. Reading a file, running a command and searching the web are all tools.
- Vibe coding
Building software by describing what you want to an AI tool and accepting what it generates, with little or no reading of the code. Fast for prototypes and personal tools; risky for anything that has to be maintained or secured.
- Workflow
A sequence of steps defined in advance. Unlike an agent, a workflow does not decide what to do next; many reliable systems combine the two, with an agent handling the steps that need judgement.