JournalE.03September 30, 202613 min read

Traces, not vibes: observability for LLM pipelines

Spans, token budgets and evals wired into the same telemetry as the rest of your stack.

d

Dipesh Chaulagain

AI-Native Fullstack Developer

OpenTelemetryEvals

Vibes don't page anyone

Most LLM features are monitored by vibes. Someone pastes a prompt into a playground, reads the answer, nods, and ships. When a customer complains a week later, the investigation is a Slack thread of screenshots and a console.log of the prompt that may or may not match what ran in production.

Meanwhile the rest of the stack has had real answers for a decade. Every HTTP request, query and queue hop is a span in a trace, sampled and indexed and alertable. The LLM call is usually the slowest, most expensive and least predictable span in the request, and it's the one we leave out.

This essay puts it back. We'll instrument a retrieval-augmented support bot with OpenTelemetry, so that model calls, tool calls, token spend and eval scores land in the same traces as your database queries. Then we'll use those traces for what vibes can't do: budgets, sampling and regression alerts.

Anatomy of an LLM trace

Here's one request to the support bot. A customer asks why an order was charged twice. The server embeds the question, searches a vector index, calls the model, lets it call a lookupOrder tool, runs a PII guardrail on the answer and scores it for groundedness. Pick a scenario and inspect the spans.

Fig. 01One request, every span
Real OpenTelemetry spans

Run a request to capture its trace.

Read the healthy trace top to bottom and the shape of an agent falls out. invoke_agent wraps the whole generateText call. Each step is one trip around the tool loop. Inside a step, chat is the actual provider call and execute_tool is your code. The two chat spans are the model deciding to call the tool, then reading the result and answering.

Now switch to Slow provider. Nothing failed, but the waterfall tells you instantly that retrieval took 130 ms and the two model calls took almost four seconds. Without spans, that's an argument about whether “the database is slow again.”

Speak gen_ai

Click a chat span and look at the attribute names: gen_ai.operation.name, gen_ai.provider.name, gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.response.finish_reasons. None of those were invented for this page. They're OpenTelemetry's GenAI semantic conventions, and they're the reason this is worth doing through OpenTelemetry rather than a bespoke LLM logging tool.

Conventions are what make telemetry portable. Honeycomb, Grafana, Datadog and the LLM-specific backends all recognise gen_ai.usage.input_tokens, so a “tokens by model” chart is one query in any of them. The conventions even cover evaluations: gen_ai.evaluation.name and gen_ai.evaluation.score.value are what the eval span above uses. Name things the standard way and every tool you'll ever adopt already understands your data.

Wire it once

Setup is one file that runs before your server handles traffic. It registers a tracer provider that batches spans to an OTLP endpoint (your collector, or a vendor directly), and installs an async context manager so spans started deep inside the AI SDK find their parent.

instrumentation.ts
import { context, trace } from "@opentelemetry/api"
import { AsyncLocalStorageContextManager } from "@opentelemetry/context-async-hooks"
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http"
import { resourceFromAttributes } from "@opentelemetry/resources"
import {
  BasicTracerProvider,
  BatchSpanProcessor,
} from "@opentelemetry/sdk-trace-base"
import { OpenTelemetry } from "@ai-sdk/otel"
import { registerTelemetry } from "ai"

context.setGlobalContextManager(new AsyncLocalStorageContextManager().enable())

const provider = new BasicTracerProvider({
  resource: resourceFromAttributes({
    "service.name": "support-bot",
    "service.version": process.env.GIT_SHA ?? "dev",
    "deployment.environment.name": process.env.NODE_ENV ?? "development",
  }),
  spanProcessors: [
    new BatchSpanProcessor(
      new OTLPTraceExporter({ url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT })
    ),
  ],
})

trace.setGlobalTracerProvider(provider)

registerTelemetry(new OpenTelemetry())

In AI SDK v7, telemetry moved out of core. @ai-sdk/otel is the integration, and registerTelemetry attaches it to every call in the process. After that, a call only needs a functionId, which becomes gen_ai.agent.name and is how you'll group these spans later. Wrap your own steps in spans with the plain OpenTelemetry API, and the model calls nest underneath them automatically:

answer.ts
import { SpanKind, trace } from "@opentelemetry/api"
import type { LanguageModel, ToolSet } from "ai"
import { generateText, isStepCount } from "ai"
import { groundedness } from "@/lib/traces/evals"

declare const model: LanguageModel
declare const tools: ToolSet
declare const search: (q: string) => Promise<Array<{ text: string }>>

const tracer = trace.getTracer("support-bot")

export const answer = (question: string) =>
  tracer.startActiveSpan(
    "answer",
    {
      kind: SpanKind.INTERNAL,
      attributes: { "app.prompt.version": "support-v7" },
    },
    async (span) => {
      try {
        const docs = await tracer.startActiveSpan("retrieve", async (s) => {
          const hits = await search(question)
          s.setAttribute("retrieval.top_k", hits.length)
          s.end()
          return hits
        })

        const result = await generateText({
          model,
          tools,
          system: "Answer from the sources. Cite them as [n].",
          prompt: `${docs.map((d, i) => `[${i + 1}] ${d.text}`).join("\n")}\n\n${question}`,
          stopWhen: isStepCount(3),
          telemetry: { functionId: "support.answer", recordInputs: false },
        })

        const g = groundedness(result.text, docs.length)
        span.setAttributes({
          "gen_ai.evaluation.name": "groundedness",
          "gen_ai.evaluation.score.value": g.score,
          "gen_ai.evaluation.score.label": g.label,
        })
        return result.text
      } finally {
        span.end()
      }
    }
  )

That nesting is the entire point. The chat span's parent is your answer span, whose parent is the HTTP span your framework created, whose parent may be a span from the browser, carried over by a traceparent header. One trace ID from the click to the token. Not a separate LLM dashboard you correlate by timestamp.

Failures that return 200

Switch Fig. 01 to Tool timeout. The orders database times out after 1.5 seconds, the AI SDK records the exception on the execute_tool span and marks it ERROR, and then it does exactly what it's designed to do: hands the error to the model, which apologises and answers from the knowledge base. The request returns 200.

That's the defining failure mode of LLM systems. The model is a very good error handler, so failures turn into plausible, worse answers instead of exceptions. Your HTTP error rate stays flat while a dependency is on fire. Alert on span status across the trace, not on response codes, and this outage shows up as the spike it is.

Look at the eval score for the same trace too: 0.5. The fallback answer is honest but only half-grounded. A timeout two layers down became a quality regression at the top, and the trace is the only artefact that connects the two.

Token budgets are SLOs

Switch to Context bloat. Someone raised retrieval from six chunks to thirty-two “to improve recall.” Input tokens quadruple, cost quadruples, latency doubles, and the root span carries a budget.exceeded event. The answer is identical.

Tokens are the one resource an LLM pipeline spends on every request, so treat them like latency: a budget per route, a distribution, and a percentile you alert on. The averages lie here as they do everywhere. Play with the knobs below and watch the tail, not the middle.

Fig. 02Where the tokens go
400 simulated requests
8
12
8.0k
budget 8.0k
0input tokens per request32k
System 0.8kHistory 2.1kQuestion 0.1kRetrieved 3.4kOutput 0.2k
p50 input
6.4k
p95 input
8.5k
Over budget
14.8%
Per 1k requests
$23.08

Every bar right of the line is a request you paid for and probably didn't need.

Two things usually surprise people. Retrieval is rarely the biggest slice for long conversations; history is, and it grows without anyone deciding it should. And trimming retrieval can't save a request whose history alone is over budget. Push history turns up with trimming on and the tail stays red. Budgets need a policy for every input, not just the one you thought of.

Enforce the budget in code, record what you dropped as span events, and feed the standard gen_ai.client.token.usage histogram so the percentile is one query away:

budget.ts
import { metrics, trace } from "@opentelemetry/api"
import { fitToBudget } from "@/lib/traces/budget"
import type { Chunk } from "@/lib/traces/budget"

const meter = metrics.getMeter("support-bot")

const tokenUsage = meter.createHistogram("gen_ai.client.token.usage", {
  unit: "{token}",
  advice: {
    explicitBucketBoundaries: [1024, 2048, 4096, 8192, 16384, 32768, 65536],
  },
})

export const withinBudget = (
  chunks: ReadonlyArray<Chunk>,
  {
    budget,
    reserved,
    route,
  }: { budget: number; reserved: number; route: string }
) => {
  const fit = fitToBudget(chunks, { budget, reserved })
  const span = trace.getActiveSpan()

  if (fit.dropped > 0) {
    span?.addEvent("budget.trimmed", {
      "budget.tokens": budget,
      "budget.dropped_chunks": fit.dropped,
    })
  }
  if (fit.tokens > budget) {
    span?.addEvent("budget.exceeded", { "budget.tokens": budget })
  }

  tokenUsage.record(fit.tokens, {
    "gen_ai.token.type": "input",
    "gen_ai.operation.name": "chat",
    "http.route": route,
  })
  return fit.kept
}

Sample the tail

A busy service can't afford to store every trace, so you sample. The default is head sampling: when a trace starts, hash its ID and keep, say, 10%. It's cheap and consistent across services, and it decides before anything interesting has happened.

Fig. 03Keep the traces that matter
2,000 traces · real OTel sampler
10%
Stored
9.6%
Errors
0/16
Slow > 4s
5/41
Over budget
2/21
Ungrounded
2/37

With head sampling at 10%, you keep 10% of everything, including about 10% of the errors, the slow requests, the budget blowouts and the hallucinations. The traces you'll actually open during an incident are exactly the ones you threw away. Raise the ratio and you pay to store thousands of boring successes to catch a few more.

Tail sampling waits until the trace is complete, then decides. Keep everything with an error, everything slow, everything over budget, everything that failed an eval, plus a small random baseline for comparison. Switch the figure to tail: storage drops and every interesting trace survives. In practice this lives in the OpenTelemetry Collector:

otel-collector.yaml
processors:
  tail_sampling:
    decision_wait: 10s
    policies:
      - name: errors
        type: status_code
        status_code: { status_codes: [ERROR] }
      - name: slow
        type: latency
        latency: { threshold_ms: 4000 }
      - name: over-budget
        type: numeric_attribute
        numeric_attribute:
          key: gen_ai.usage.input_tokens
          min_value: 8001
      - name: ungrounded
        type: ottl_condition
        ottl_condition:
          error_mode: ignore
          span:
            - attributes["gen_ai.evaluation.score.value"] < 0.5
      - name: baseline
        type: probabilistic
        probabilistic: { sampling_percentage: 5 }

The rules only work because the LLM data is on the spans. You can't tail-sample on groundedness if groundedness lives in a spreadsheet. One operational catch: every span of a trace has to reach the same collector instance, so put a load balancer that routes by trace ID in front of a collector fleet.

Evals are telemetry

Offline evals, a golden set scored in CI, catch regressions you anticipated. Production catches the rest, and only if you score real traffic. Cheap checks like groundedness, format validity or refusal detection can run inline on every request. Expensive LLM-as-judge scores run on a sample, after the response has gone out.

Either way, write the score onto a span. Then an eval score is just another attribute you can group, filter and alert on, and the most useful thing to group it by is the prompt version:

Fig. 04A regression hiding in the average
Groundedness, 24h, simulated
10%
alert < 0.8deploy support-v800:0006:0012:0018:0023:45
all traffic
No alert. The regression is invisible.

At a 10% canary, the new prompt's drop in groundedness barely moves the global average. No alert fires; the regression reaches 100% of traffic next week. Group by app.prompt.version and the canary line falls through the threshold within fifteen minutes of the deploy. Same data, one attribute. Put the version of everything (prompt, model, retrieval index) on the span.

When the scoring runs later, in a queue or a batch job, start a new trace and link it to the original span, instead of pretending it's a child of a request that ended minutes ago. Backends render the link, so you can jump from a bad score straight to the request that earned it:

eval-later.ts
import { trace } from "@opentelemetry/api"
import type { SpanContext } from "@opentelemetry/api"
import { groundedness } from "@/lib/traces/evals"

const tracer = trace.getTracer("support-bot.evals")

export const scoreLater = (
  answer: string,
  sources: number,
  origin: SpanContext
) =>
  tracer.startActiveSpan(
    "evaluate groundedness",
    { root: true, links: [{ context: origin }] },
    (span) => {
      const g = groundedness(answer, sources)
      span.setAttributes({
        "gen_ai.operation.name": "evaluate",
        "gen_ai.evaluation.name": "groundedness",
        "gen_ai.evaluation.score.value": g.score,
        "gen_ai.evaluation.score.label": g.label,
      })
      span.end()
      return g
    }
  )

Prompts are PII

The fastest way to turn an observability project into a security incident is to record full prompts and completions by default. Prompts contain whatever users type: names, addresses, order numbers, the occasional password. Retrieved chunks contain your internal documents. Traces get copied into vendors, retained for months and read by people who'd never be granted database access.

So the pipeline here sets recordInputs: false and recordOutputs: false, and records the shape instead: token counts, chunk counts, scores, finish reasons, versions. That covers nearly every question you'll ask of a trace. When you do need content, capture it for a sampled slice, redact in the collector before it leaves your network, and give it a shorter retention than your metrics.

The same goes for exceptions. The tool error in Fig. 01 had a stack trace full of file paths; the route strips exception.stacktrace before returning spans to the browser. Know what your spans carry before you ship them anywhere.

The checklist

  • LLM calls export through OpenTelemetry, in the same traces as the rest of the request.
  • The AI SDK integration emits gen_ai.* attributes; you don't hand-roll them.
  • Every call has a functionId; every span has prompt, model and index versions.
  • Alerts use span status across the trace, not HTTP status codes.
  • Each route has a token budget, enforced in code and tracked at p95.
  • Tail sampling keeps errors, slow, over-budget and failed-eval traces.
  • Eval scores are span attributes; late evals link to their origin span.
  • Prompt and completion capture is off by default and redacted when on.

None of this is new. It's the observability playbook we already run for databases and queues, applied to the component that most needs it. Instrument the model like any other dependency and you get evidence to argue with, not vibes. For keeping the pipeline alive once you can see it failing, see E.01.