Not all OpenTelemetry traces are created equal. ⚡
If you're instrumenting an AI agent, the semantic convention you choose shapes how much you'll actually be able to debug later.
When we built the Arthur engine, we evaluated both options and went with OpenInference over the OTEL-community GenAI conventions. Here's why it matters for production agents:
→ Richer LLM span detail — full prompts, completions, token counts, cost, and model parameters
→ First-class retrieval and re-ranking spans, which RAG-heavy agents live and die by
→ Clear span typing — LLM, TOOL, AGENT, CHAIN, RETRIEVER are all distinct, not lumped together
→ Explicit message, document, and tool-call types you can actually query against
The OTEL GenAI conventions are improving, but compare two traces side by side and the expressiveness gap is obvious (Images: OTEL GenAI semantic on left and OpenInference semantic on right).
This is the kind of decision that feels minor on day one and compounds for months.
Part 1 of our Best Practices for Building Agents series goes deeper on what to trace, which frameworks ship with strong OTEL support out of the box, and how our FDE team approaches observability with enterprise customers.
Read it here (link in comments) 👇
ALT A trace using OTEL GenAI semantic conventions
ALT A trace using OpenInference semantic conventions