# Enterprise AI Glossary: 40+ Terms, Plainly Explained

A plain-language glossary of 40+ enterprise AI terms — RAG, agents, evals, guardrails, fine-tuning, and more — for technical and business readers.

This glossary defines the terms most often used in enterprise AI work, from large language models and retrieval pipelines to agents, evaluation, and governance. Definitions are written for mixed technical and business readers and assume no prior machine learning background. For a broader view of how these concepts fit into a delivery program, see our [capabilities](/capabilities/) and [engagement models](/engagements/).

## Foundations

### LLM (large language model)

A statistical model trained on very large amounts of text so it can predict, complete, and generate language. Modern LLMs can summarize documents, draft reports, answer questions, and write code when given clear instructions. They are probabilistic systems, not databases, so their outputs should be verified for high-stakes uses.

### Foundation model

A large model trained on broad data that can be adapted to many different tasks without being rebuilt from scratch. LLMs such as the GPT, Claude, and Llama families are foundation models for language. Enterprises typically build on a foundation model rather than training their own, then adapt it with retrieval, fine-tuning, or instructions.

### Tokens

The small units of text a model reads and writes, roughly equivalent to word fragments. Models charge, measure, and limit work in tokens rather than words or pages. Understanding token counts matters for budgeting and for staying within a model's context window.

### Context window

The maximum amount of text a model can consider in a single request, measured in tokens and including both the input and the output. A larger window lets a model reason over longer documents or conversations at once. Content beyond the window is simply not seen by the model, so inputs must be selected and condensed deliberately.

### Embeddings

Numeric representations of text that capture meaning, allowing a computer to judge how similar two pieces of content are. They power search over documents, recommendations, and clustering by turning language into coordinates in a mathematical space. Most retrieval systems depend on a separate embedding model, distinct from the LLM that generates answers.

### Fine-tuning

Additional training on a focused dataset to adapt a foundation model to a particular domain, tone, or task. It can improve accuracy on specialized vocabulary and consistent formatting, but it does not reliably teach a model new facts — retrieval is usually the better mechanism for knowledge. Fine-tuned models still require evaluation, since behavior can change in unexpected ways.

### Prompt engineering

The practice of designing the instructions and examples given to a model to get reliable, well-formatted output. It includes writing system prompts, structuring inputs, and iterating on wording based on observed results. It is the lowest-cost way to improve an application, but it rarely substitutes for proper evaluation and guardrails.

### System prompt

The set of instructions given to a model at the start of a conversation that define its role, constraints, and behavior. It sits behind the user's input and shapes every response the model produces. Well-written system prompts are one of the simplest controls for keeping an assistant consistent and on-policy.

### Few-shot prompting

Providing a handful of worked examples inside the prompt so the model can imitate the pattern. It is often enough to teach formatting, classification rules, or extraction schemas without any training. The examples consume context window space, so they are best kept short and representative.

### Temperature

A setting that controls how random a model's output is. Low values make responses more deterministic and repeatable; higher values make them more varied and creative. Most enterprise applications run at low temperatures to keep answers consistent.

### Hallucinations

Confident-sounding but false statements a model generates, including fabricated facts, citations, or references. They are a natural consequence of how models predict text and cannot be eliminated entirely. Grounding answers in retrieved sources and requiring citations are the standard ways to manage them.

## Retrieval and RAG

### RAG (retrieval-augmented generation)

An architecture in which the model first retrieves relevant passages from a knowledge base, then uses them to compose its answer. It lets a general-purpose model answer questions about private or current organizational information it was never trained on. Most enterprise AI deployments use RAG because it is cheaper than fine-tuning and keeps answers tied to verifiable sources. See our guide on taking [RAG from pilot to production](/guides/rag-from-pilot-to-production/) for the full delivery picture.

### Chunking

Splitting source documents into smaller passages before indexing them for retrieval. Chunk size and overlap determine what the model sees as context: too large and answers get vague, too small and they lose necessary context. Chunking strategy is one of the highest-leverage decisions in a RAG pipeline.

### Vector database

A storage system built to find content by meaning rather than by keywords, using embeddings. When a user asks a question, the database returns the passages whose embeddings are closest to the question's embedding. Common enterprise choices include purpose-built vector stores and vector extensions to existing databases.

### Hybrid search

Combining keyword search with embedding-based search so that both exact terms and semantic meaning influence retrieval. It performs better than either method alone, especially for jargon, acronyms, and product names where pure semantic search can miss. Most production retrieval systems use some form of hybrid ranking.

### Reranking

A second, more careful relevance check applied to the initial set of retrieved passages before they reach the model. A reranker scores each candidate against the question and keeps only the strongest matches, improving answer quality at a modest latency cost. It is a standard addition once a pilot moves toward production.

### Grounding

Restricting the model's answer to information found in the retrieved sources rather than its own training data. Grounded systems answer only from approved content and decline or flag questions they cannot support. It is the primary design pattern for trustworthy enterprise question-answering.

### Citations

References to the specific source passages the model used, shown alongside its answer. They let a reader verify claims and give auditors a trace from output back to input. Requiring citations is both a usability feature and a control.

## Agents

### AI agent

A system in which a model is given tools and a goal, and decides for itself which actions to take and in what order. Unlike a single question-and-answer call, an agent works through multiple steps, using intermediate results to guide its next move. Enterprises use agents for research, triage, and multi-step workflows that would otherwise require manual coordination.

### Tool use / function calling

The mechanism by which a model triggers external actions, such as querying a database, calling an API, or sending a message. The model outputs a structured request describing which tool to call and with what arguments; the application executes it and returns the result. Reliable tool definitions and strict argument validation are what make agents safe to operate.

### Planning

An agent's ability to break a goal into a sequence of steps and adjust the plan as new information arrives. Simple agents follow fixed workflows; more capable ones reason about dependencies, alternatives, and error recovery. Planning quality is one of the main factors separating prototype agents from dependable ones.

### Multi-agent systems

Designs in which several agents with distinct roles collaborate on a task — for example, one that drafts, one that critiques, and one that checks facts. Specialization can raise quality on complex work, but it also multiplies cost, latency, and failure modes. They are best reserved for tasks where the gains have been demonstrated in evaluation.

### Reasoning trace

The step-by-step internal record an agent or model produces while working toward an answer. Reviewing traces helps engineers understand why an agent took a particular action and where it went wrong. Traces can contain sensitive or misleading intermediate thoughts, so they are typically shown to operators, not end users.

### Human-in-the-loop

A design in which a person reviews and approves an agent's proposed actions before they take effect. It is the standard control for consequential actions such as sending customer messages, changing records, or spending money. The approval points, escalation paths, and override audit records all need to be designed deliberately.

## Evaluation and safety

### Evals

Systematic tests that measure whether an AI application behaves as required, using curated examples, automated checks, and human review. They cover accuracy, format compliance, tone, refusal behavior, and regressions after each change. Running evals before and after every update is what makes AI development disciplined rather than hopeful. See our guide on [evaluating LLM applications](/guides/evaluating-llm-applications/) for a practical method.

### LLM-as-judge

Using one model to score or compare the outputs of another, according to written criteria. It scales far beyond manual review and is useful for ranking drafts or catching obvious failures. Because judges have their own biases, their scores should be calibrated against human judgment before being trusted.

### Red-teaming

Deliberately probing a system with adversarial inputs to find ways it can be misled, bypassed, or made to misbehave. It covers prompt attacks, misuse scenarios, and edge cases the designers did not anticipate. Findings feed back into guardrails, system prompts, and monitoring.

### Guardrails

Rules and filters placed around a model to enforce policy: blocking disallowed content, masking sensitive data, and constraining outputs to approved formats and topics. They operate at input, output, and tool-use layers. Guardrails reduce risk but are not perfect, so they complement rather than replace testing.

### Prompt injection

An attack in which malicious instructions hidden in data the model reads — a document, web page, or tool result — attempt to override the system's instructions. It is one of the most discussed vulnerabilities in deployed AI systems. Defenses include separating instructions from data, validating tool inputs, and requiring human approval for sensitive actions.

### Jailbreak

Attempts to get a model to violate its own safety policies, often through crafted roleplay, hypotheticals, or encoded instructions. Vendors patch known jailbreaks continuously, and new ones surface regularly. Application-level guardrails and monitoring provide a second line of defense.

### Alignment

The research and engineering effort to make models behave in line with human intent and values: helpful, honest, and safe. It covers training techniques, evaluations, and deployment safeguards. For enterprise buyers, a vendor's alignment work shows up as model behavior, safety documentation, and refusal quality.

## Deployment and operations

### Inference

Running a trained model to produce outputs in response to inputs. It is the operational phase of AI: what happens each time a user asks a question or an agent takes a step. Inference cost, speed, and reliability are the core operational concerns for any deployed system.

### Latency / TTFT

The delay between a request and a response; TTFT (time to first token) measures how quickly the first part of a streamed answer arrives. Streaming output as it is generated makes long answers feel fast even when total time is unchanged. Latency budgets shape model choice, architecture, and user experience design.

### Throughput

The volume of requests or tokens a deployment can handle per unit of time. Planning throughput means sizing infrastructure for peak load, not average load. For hosted APIs, throughput is governed by rate limits and pricing tiers.

### Quantization

Reducing the numerical precision of a model's weights so it runs faster and uses less memory, with only a small loss in quality. It makes it practical to run capable models on modest hardware or to cut hosting costs. Quantized models still need evaluation, since quality effects vary by task.

### Distillation

Training a smaller model to imitate the behavior of a larger, more capable one. The result is a cheaper, faster model tuned for a specific task. It is a common route to production economics when a flagship model proves the use case but is too expensive to run at scale.

### Observability for AI

Logging and monitoring everything needed to understand an AI system in production: inputs, retrieved context, model outputs, tool calls, latency, cost, and user feedback. Because model behavior is probabilistic, observability is how teams detect drift, catch failures, and gather data for the next round of evals. It is a prerequisite for operating AI responsibly at scale.

### CI/CD for ML

Applying continuous integration and delivery practices to AI systems: versioned prompts, models, and datasets, automated evals on every change, and staged rollouts with rollback. Traditional software CI/CD is not enough, because changes to data or prompts can alter behavior without touching code. Mature teams treat prompts and evals as first-class artifacts in their pipelines.

## Governance

### AI governance

The set of policies, roles, and processes an organization uses to oversee its AI systems: who may build and deploy them, what risk controls apply, and how compliance is demonstrated. It connects technical controls to accountability, from model selection through retirement. Regulated sectors and [government and enterprise](/government-enterprise/) buyers increasingly require it as a condition of adoption.

### Risk tiers

Classifying AI use cases by the harm they could cause, so that controls scale with risk. A drafting assistant faces lighter requirements than a system that screens benefit applications. Tiering focuses effort where it matters and is a central idea in emerging AI regulation. See our [practical guide to AI governance](/guides/ai-governance-practical-guide/) for how organizations operationalize it.

### Model inventory

A register of every AI model in use across the organization: what it is, where it came from, what it is used for, and who owns it. It is the starting point for any governance program, since you cannot govern what you have not catalogued. Inventories typically capture version, data sources, risk tier, and approval status.

### Audit trail

A complete, tamper-evident record of an AI system's activity: prompts, retrieved sources, model outputs, tool calls, human approvals, and timestamps. It supports incident investigation, regulatory examination, and customer disputes. Building audit logging in from the start is far cheaper than reconstructing it later.

### Data lineage

Documentation of where the data in an AI system comes from, how it is transformed, and where it ends up. It answers questions about consent, licensing, and freshness, and it is essential when retrieved content feeds answers to users. Lineage becomes a procurement and legal concern as soon as third-party or personal data is involved.

### Responsible AI

The organizational commitment to develop and deploy AI in ways that are fair, transparent, accountable, and respectful of privacy and safety. It is expressed through principles, training, review boards, and the controls described throughout this glossary. Done well, it is a trust asset with customers, regulators, and staff.

If these concepts are clear but the path from pilot to production is not, the right next step is a structured conversation about scope, risk, and delivery. Review our [engagement models](/engagements/) or [contact us](/contact/) to discuss how these ideas apply to your organization.
