# Observability for AI Systems: Tracing, Evals, and Cost Telemetry

What observability means for AI systems: LLM tracing, evaluation in production, quality signals, and cost telemetry for every request.

Conventional observability was built for deterministic systems. Microservices return the same output for the same input, fail in repeatable ways, and produce metrics that mean something stable over time. AI systems break these assumptions. A RAG assistant can answer the same question well on Tuesday and poorly on Thursday because a document was updated, a reranker threshold shifted, or the model provider released a new minor version. Token costs vary with every response. Quality is a judgment, not a status code. Microservice dashboards — latency percentiles, error rates, throughput — still matter, but they miss the failure modes that sink AI deployments: confident wrong answers, silently degrading retrieval, costs drifting upward without matching gains in quality.

Observability for AI systems means instrumenting those failure modes directly: tracing each request through retrieval, prompting, and tool calls; scoring quality on live traffic rather than only in pre-release benchmarks; accounting for cost per request the way finance accounts for spend per transaction; and setting alerting and service-level objectives that reflect what the system promises its users. This guide walks through each layer and closes with how to make telemetry safe for production data.

## Tracing LLM requests

A single AI response is rarely a single model call. A typical RAG request fans out into query rewriting, hybrid search, reranking, prompt assembly, model calls, and often tool calls — each step with its own inputs, outputs, latency, and cost. None of that intermediate state is visible in a standard application log, so a bad answer cannot be diagnosed from it. Distributed tracing adapted for LLM workflows recovers what the log loses.

### Spans for retrieval, prompts, and tool calls

Structure every request as a trace of spans. At minimum, instrument these:

- **Request span**: the request envelope — session or conversation identifier, user or tenant identifier, the feature or use case that triggered the request, timing, final outcome.
- **Retrieval spans**: the query issued to each retrieval source (vector search, keyword search, metadata filter), the candidate counts, the reranker scores, and the identifiers of the chunks ultimately passed forward. A retrieval span turns "the answer was wrong" into "the answer was wrong because the top chunks came from an outdated document version."
- **Prompt spans**: the assembled prompt, the model and version called, the parameters used (temperature, max tokens, stop sequences), and the input and output token counts.
- **Tool call spans**: for agentic systems, each tool invocation — the tool name, its arguments, its result or error, and its latency. Tool spans reveal when the model is fine but a dependency it called is not.

Keep the span model aligned with existing standards. OpenTelemetry's span conventions provide correlation, sampling, and export, and most managed observability platforms now understand LLM-specific span attributes — keeping AI traces in the same system as the rest of the application's telemetry rather than in a sidecar dashboard nobody opens.

### What to capture per span

The temptation is to record everything; the discipline is to record what changes a decision. Per span, capture:

- **Inputs and outputs**, or references to them. Full prompts and completions are valuable but expensive to store and risky for privacy — store hashes or truncated versions by default, with full capture behind a sampling flag for debugging.
- **Identifiers**: request, session, tenant, and user ids (pseudonymized where appropriate), model name and version, corpus or index version.
- **Measurements**: start and end timestamps, token counts split by input and output, cost in the billing currency, retry count.
- **Evaluative metadata**: prompt template version, retrieval configuration, which guardrails ran and whether any fired. Without this, quality changes cannot be attributed to specific deployments.

Recording every span at full fidelity is simple but costly, so adopt a deliberate sampling policy: sample at a fixed rate, and keep 100 percent of error traces, traces where a guardrail fired, and traces that received negative user feedback. Storage stays bounded, and the traces you most need — the failures — are always available.

## Quality signals in production

Offline evaluation benchmarks are necessary and insufficient: they measure a frozen dataset against a frozen model, while production is a moving system — documents change, user queries drift, and model providers update their models without asking. Quality signals in production close that gap by scoring live traffic continuously.

### Eval scores on live traffic

Run a lightweight evaluation layer over a sample of production requests — no human graders on every answer required. Common patterns include:

- **LLM-as-judge scoring**: a second model scores the response for correctness, groundedness in the retrieved context, and format adherence. Calibrate the judge against human labels on a representative sample before trusting its scores, and re-calibrate whenever the model or the task changes.
- **Reference checks**: for questions with known answers, compare the response against expected content or required citations — a hard correctness metric for a defined slice of traffic.
- **Behavioral checks**: deterministic assertions that require no model at all — did the response cite the sources it was given, refuse when the retrieval set was empty, stay within the allowed language, emit valid structured output. Cheap to run on every request, and they catch entire classes of failure.

Store eval scores as attributes on the request trace so quality, cost, and latency can be analyzed together. The useful production question is rarely "what is our average score" — it is "which retrieval configurations, prompt versions, and document subsets produce low scores," which requires the trace context.

### User feedback, including thumbs-down analysis

Explicit feedback — thumbs up or down, ratings, corrected answers — is the highest-value signal available, and most teams underuse it. Collect it in the product, and treat the analysis as an engineering function, not a dashboard decoration.

Cluster negative feedback by topic and failure type — wrong answer, incomplete answer, refused to answer, hallucinated citation, slow response, formatting problem. For each cluster, pull the traces and determine the root cause: retrieval (wrong documents), generation (wrong use of right documents), or a coverage gap (the corpus does not cover this topic). Each has a different fix — the index, the prompt or model choice, or the content — which is why the taxonomy matters.

One caution: feedback is biased — dissatisfied users rate more often than satisfied ones, and silence is not approval. Use it to find and prioritize failures, not to compute an absolute quality score.

## Cost telemetry

AI requests have a unit economics problem that traditional requests do not: a web request costs roughly the same every time, while an AI request's cost depends on prompt length, retrieval and reranking steps, model choice, tool calls, and output length. Costs compound silently — a slightly more verbose prompt template, rolled out quietly, can raise monthly spend substantially before anyone notices.

### Per-request token accounting

Price each request from the provider's published pricing for the exact model and version called, stored as configuration. For self-hosted models, use an internal unit cost per token or per GPU-second. The point is that every request carries a dollar figure, not that the figure is perfect.

Break token accounting down by step. If query rewriting and reranking consume 40 percent of a request's tokens, a smaller rewriting model or a tighter retrieval budget may hold quality while cutting cost. Without per-step accounting, only the final model call is visible.

### Cost attribution by team and use case

Aggregate cost along the dimensions the organization budgets against: product or feature, team, tenant or customer, environment, and time. Attribution turns "the AI bill went up" into "the support assistant's cost per resolved conversation rose after the new prompt template shipped" — the framing a team needs to decide whether the quality gain justified the spend.

Pair cost dashboards with quality dashboards on the same screen. Cost without quality context produces a race to the cheapest acceptable output; quality without cost context produces prompts nobody can afford to scale. The operating metric that matters for most enterprise AI products is cost per successful outcome — per resolved ticket, per correctly answered question — not cost per token.

## Latency and reliability

### TTFT, streaming, and perceived latency

AI responses stream token by token, so time-to-first-token (TTFT) matters as much as total duration. A request that takes eight seconds but streams from second one feels fast; a request that takes four seconds with nothing until the end feels broken. Instrument both.

Track TTFT, inter-token gaps, and total completion time separately. A slow TTFT usually points at retrieval, queueing, or prompt assembly; growing inter-token gaps point at serving or network; a slow total time with healthy streaming is usually just a long response. A single "latency" metric conflates these and misdirects investigation.

Streaming complicates failure handling: a response that fails halfway has already shown the user partial content, and retrying from scratch is jarring. Design the client to handle mid-stream failures explicitly — either by completing the visible partial response gracefully or by restarting with a clear indicator — and trace both the failure and the recovery path.

### Retries, timeouts, and fallback behavior

Treat provider calls like any unreliable dependency: bounded retries with backoff and jitter, explicit timeouts per call type, and circuit breaking on error-rate spikes. AI-specific is the fallback design: when the primary model is unavailable or rate-limited, options include a smaller or cheaper model with a narrower capability promise, a cached or template response for known queries, or a graceful degradation that tells the user the system is limited rather than producing a low-quality answer silently.

Log every fallback invocation: which fallback engaged, why, and what it cost in quality and latency. Fallbacks that engage quietly and often are a reliability problem disguised as a solution — review fallback engagement rates the way you review error rates.

## Alerting and SLOs for AI

SLOs for AI systems must cover what the system actually promises: a useful, correct-enough, timely answer at a sustainable cost. Uptime alone is hollow for a system whose characteristic failure is a confident wrong answer from a healthy endpoint.

### What an SLO looks like for a RAG assistant

A workable SLO set for an internal RAG assistant might read:

- **Availability**: 99.5 percent of requests return a completed response (not a provider error, timeout, or unhandled exception) over a 30-day window.
- **Latency**: 95th-percentile TTFT under 3 seconds and 95th-percentile total duration under 20 seconds, measured on streaming responses.
- **Groundedness**: at least 98 percent of factual claims cite retrieved sources, measured by automated checks on sampled traffic.
- **Eval quality**: mean LLM-judge correctness of at least 4.0 out of 5 on the weekly production sample, with no use case below 3.5.
- **Cost**: cost per completed request within 15 percent of the budgeted baseline for the quarter.

The numbers are illustrative; each organization sets its own. The structure is the point: availability, latency, quality, and cost, each with a defined measurement and window. Note what is absent — no SLO on "never wrong," which no AI system can promise. SLOs describe the behavior envelope, including the acknowledged error rate.

### Alerting on the right signals

Alert on leading indicators and SLO burn, not on every anomaly. Practical alerting tiers:

- **Page-level**: provider outage or sustained error-rate spike, SLO burn rate exceeding the fast-burn threshold, guardrail firing rates that suggest an attack or a data problem.
- **Ticket-level**: slow-burn SLO degradation over days, a step change in cost per request after a deployment, eval scores drifting downward week over week, a new cluster of thumbs-down feedback forming around a topic.
- **Review-level**: weekly quality and cost reviews driven by the dashboards, not by alerts. This is where gradual drift gets caught — retrieval quality decaying as documents age, prompt templates accreting complexity, cost per outcome creeping up.

An AI system produces far more candidate signals than a conventional service; if everything pages, nothing does. Keep page-level alerts to the small set that means users are harmed right now, and route everything else to tickets and reviews.

## Privacy in telemetry

AI telemetry is unusually privacy-sensitive: traces contain user questions, retrieved documents, and model outputs, any of which may include personal information, pasted credentials, or confidential content. A log-first pipeline creates a second copy of sensitive data in a system with its own access model and retention policy.

### Redacting PII from traces

Redact at capture time, in the application, before spans leave the process. Handle the common patterns: email addresses, phone numbers, government identifiers, payment card numbers, API keys and tokens. In free text, run a PII detection pass and replace matches with typed placeholders such as `[EMAIL]` or `[ACCOUNT_NUMBER]` rather than deleting them, so the trace remains structurally useful.

Never capture some spans in full: authentication material, health information, and content from restricted document classifications. Exclude them by policy, enforce the exclusion in code, and test that the telemetry output contains no raw values. Treat the telemetry schema as a data contract with the same rigor as an API contract.

### Retention, access, and the right to be forgotten

Keep retention explicit and short: raw prompts and completions for 30 days, redacted traces for 90 days, aggregated metrics for a year or more. Debugging rarely needs a six-month-old prompt, and every retained day is a day of exposure. Aggregates (eval scores by configuration, cost by use case, latency distributions) keep their analytical value long after the raw text should be gone.

Limit raw-trace access to the engineers who debug the system, behind authentication and audit logging; never use production data for model training without explicit consent and a documented legal basis. And make deletion work end to end: when a user or a regulator asks for data to be erased, the request has to reach the telemetry store, not just the application database. Build the deletion path before you need it.

For organizations working through how telemetry fits into their broader AI governance posture, our [capabilities](/capabilities/) outline the areas we advise on.

## Putting it into practice

Observability for AI is not a product you buy; it is instrumentation built into the system from the start, plus habits — trace-first debugging, weekly quality reviews, cost-per-outcome tracking — that the team practices continuously. A workable sequence: tracing and per-request cost accounting first — cheap, and they make every later decision evidence-based; production evals and feedback analysis second — they reveal actual quality; then SLOs and alerting on trusted signals, with redaction and retention controls wrapped around the pipeline before it accumulates liability.

None of this requires exotic tooling: OpenTelemetry traces, token accounting in the request path, a sampled eval harness, and dashboards that put quality next to cost carry most organizations further than a feature-rich platform with no operating discipline. The discipline is the hard part, and it is where outside perspective helps most.

If your team is standing up AI observability — or discovering that the dashboards built for your microservices are not answering the questions your AI system raises — an [advisory engagement](/engagements/) can help define the telemetry architecture, the SLO framework, and the operating cadence, or you can [contact us](/contact/) to discuss where your current setup stands.
