A trace is the complete, structured record of one request's journey through your AI application. It captures the user input, every intermediate step, and the final output, with timing. A span is a single operation within that trace: one LLM call, one retrieval query, one tool execution, one guardrail check.
Spans nest, so a "handle_question" root span might contain a "retrieve_docs" span and two "llm_call" spans.
Each span carries attributes: model name, prompt and completion (or references to them), input/output token counts, cost, latency, and error status. Spans also carry custom metadata like user tier or feature flag. This structure makes LLM apps debuggable.
When a user reports a bad answer, the trace shows the retrieved documents and final prompt. It also shows what the model returned at each step.
Traces are also the raw material for everything else in the eval stack. Judges score traces, datasets are curated from traces, and cost/latency dashboards aggregate span attributes. Tools like LangSmith, Langfuse, Braintrust, and Arize Phoenix are all built around this trace/span model, increasingly using OpenTelemetry conventions.
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.
Why there's no diagram: “”
The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓
The diagram below the answer is the concept . Jump to it ↓