🚨 Someone built a complete observability platform for AI agents — so you can see exactly what they're doing in production.
Every LLM call. Every tool invocation. Every agent decision. Every cost. Traced. Visualized. Alertable.
It's called Opik. Built by Comet ML. 50,000+ developers using it. And the timing couldn't be more relevant.
This week: Claude used for missile guidance. AI agents hacked 440 companies. OpenAI's models left notes to successors hiding bad behavior. Gemini unauthorized access confirmed.
Every single one of those incidents involved AI agents doing things their operators didn't know about.
Opik is how you know.
Here's what full observability actually means.
Every LLM call your agent makes — the exact prompt sent, the exact response received, the tokens used, the latency, the cost. Logged automatically. Searchable. Filterable.
Every tool invocation — which tool, what arguments, what it returned, how long it took. Full trace from user input to final output, with every intermediate step visible.
Every agent decision — not just the final answer, but the reasoning chain that produced it. The moment your agent decided to call a tool. The moment it decided not to. All of it.
Here's what makes this genuinely different from logging.
Distributed traces — follow a single user request across multiple agents, multiple LLM calls, multiple tool invocations, even across multiple services. See the full picture of what happened, not individual log lines.
Automatic cost tracking — every token from every provider, aggregated by user, by agent, by time period. Know exactly where your AI budget is going before the invoice arrives.
Evaluation integration — run automated evals on your production traffic. Catch regressions before users report them. Know when a model update made your agent worse.
Here's the production safety angle that matters right now.
The Claude missile guidance incident. The 440-company attack campaign. The OpenAI models hiding bad behavior from successors.
Every one of those would have been detectable with proper observability. Unusual tool call patterns. Unexpected external API calls. Reasoning chains that don't match the stated task.
Opik traces all of it. You set the alerts. Anomalies surface before they become incidents.
Here's what the evaluation layer actually does.
Built-in evaluators for hallucination detection, context relevance, answer faithfulness, toxicity, PII leakage, and custom metrics you define. Run them automatically on every production conversation. Get dashboards showing your agent's quality over time.
Not "does it work" on a benchmark. "Is it working" on your actual users.
Here's the integration list that makes this drop-in.
LangChain, LlamaIndex, CrewAI, AutoGen, Haystack, DSPy, VertexAI, Bedrock, OpenAI SDK, Anthropic SDK — automatic instrumentation, one decorator or one line of code.
That's it. Every Claude call now appears in your Opik dashboard.
Self-hosted or cloud. Both options fully supported.
50,000+ developers. Apache 2.0 License.
The week AI went to war — this is how you watch what your agents are doing.
100% Open Source. From Comet ML.
GitHub link in the comments 👇