Active observability for agents in production.

Your team deserves more from your agent observability platform - a single, connected place for instrumentation, investigation, and measurement, enhanced with intelligence. In Braintrust, you can start with an open-ended prompt about agent behavior and carry the investigation into code, evals, automation, and monitoring. New tools like Patterns and Debugger make it easier to find recurring behavior, get to root causes, and act on what you learn. The context you build in one step automatically follows to the next. Work in the platform with an enhanced Loop experience, or move to your coding agent of choice. Read more → braintrustdata.link/agent-ob…
13
11
56
42,021
GPT-6 Sol and Luna are now available in Loop. Bring your own provider keys, or use with Braintrust native inference. Ask Loop to investigate traces, logs, experiments, and datasets in your Braintrust projects. Or use automations to repeat open-ended investigations on a schedule. Get started → braintrustdata.link/loop-gpt…
5
471
When and how should you use Jev? We've been evaling to find out. - Using Jev in Braintrust → braintrustdata.link/jev-eval - Testing how Jev holds up as a judge → braintrustdata.link/jev-vs-g… - Scoring outputs with Jev → braintrustdata.link/jev-how-…
3
16
872
We built Brainstore for the amount of data agents generate. Now agents are helping teams search through traces and run far more queries than a person would. Nitro is Brainstore's new asynchronous query execution engine. It separates waiting for storage from processing the data it returns, bypassing the limits of traditional object storage. We ran Nitro on more than 1 million real queries and found that full-text searches were over 2x faster. It's already on for Braintrust SaaS and BYOC customers. Read more → braintrustdata.link/nitro-bl…
3
7
42
10,189
/bro and i-have-adhd are OSS skills designed to make LLM writing more concise and conversational, much like Opus 5.5 claims. But can AI de-slop itself? /bro reduced writing quality by about 22–26 percentage points and task correctness by 17 points. i-have-adhd showed no reliable improvement on writing quality, while task correctness fell by about 11 points for Opus and 3 points for Fable. Results may differ on other tasks, but we found that style instructions still need to be evaled alongside task performance. Read more → braintrustdata.link/opus-55-…
1
11
704
Everyone is tired of reading AI slop. Anthropic says Opus 5.5 writes more naturally and actually follows instructions, so we ran an eval. We tested it against the new versions of OpenAI's GPT-6 Luna and Sol to see whether models can solve a problem AND write a decent explanation. Opus 5.5 came out on top in our tests, solving 80.1% of tasks and scoring 84.4% on writing quality. But Luna got to 78.0% and 83.8% for roughly 2% of the cost. Test them on your own tasks to see whether the difference matters. Read more → braintrustdata.link/opus-55-…
4
32
3,301
In Braintrust, you can use Jev both to score outputs and trace an agent that uses Jev to make decisions. We show you how with Jevtris, a browser-based clone of everyone's favorite puzzle game. Watch the tutorial → braintrustdata.link/jev-how-…
1
7
813
We tested where Jev holds up as a judge. Jev was fast, cheap, and highly competitive for judging groundedness. But it lagged behind models with reasoning for math and code domains. Use it for certain eval tasks, but don't throw out your LLM-as-a-judge just yet. Read more → braintrustdata.link/jev-vs-g…
Jev from @typesafeai can replace your LLM-as-a-judge for scoring agent responses. It returns a choice or numeric result with information about uncertainty, so you don't need to spend time and resources prompting a general-purpose model into an LLM judge. Use Jev as a judge scorer in Braintrust and review its selected answer, confidence, and probabilities alongside the score. Trace Jev calls from your own application with the JavaScript or Python SDK. Read more → braintrustdata.link/jev-eval
10
4
38
4,513
Jev from @typesafeai can replace your LLM-as-a-judge for scoring agent responses. It returns a choice or numeric result with information about uncertainty, so you don't need to spend time and resources prompting a general-purpose model into an LLM judge. Use Jev as a judge scorer in Braintrust and review its selected answer, confidence, and probabilities alongside the score. Trace Jev calls from your own application with the JavaScript or Python SDK. Read more → braintrustdata.link/jev-eval
20
9
185
57,399
With @gauge_sh, brands can control their presence on all the major AI platforms, including chat and agent recommendations. Patterns identifies inefficiencies in their agent and helps them resolve issues automatically. Try Patterns → braintrustdata.link/active-o…
1
3
599
Tolan builds AI companions that engage in natural voice conversations and rely on complex memory systems. Patterns is how they find the root causes of agent failures. Try Patterns → braintrustdata.link/agent-ob…
5
6
641
Braintrust retweeted
tried out this skill I found that uses statistical best practices to analyze my AI eval experiments
1
1
2
302
Braintrust retweeted
traces are easy to read when you have 1-liner for each span
1
2
6
620
What's new: - Surface hidden insights in your traces with Patterns - Schedule custom recurring analysis of your traces with Loop automations - Keep multiple human reviews independent with blind reviews - Manage access to your org-level AI providers with fine-grained permissions - Score your multi-turn traces as one unit with group-scoped scores - Search, star, and section your dashboards Read more → braintrustdata.link/whats-ne…
6
919
When you eval open-weight models, you also need to eval the serving stack. The inference engine, model revision, precision, caching, and account limits can all impact the system you're testing. We ran an eval on Kimi K3 served on @FireworksAI_HQ vs @Kimi_Moonshot. They performed equally well on quality, but median time to first token was 55% faster on Fireworks. Read more → braintrustdata.link/moonshot…
1
2
8
984
Engineers at @evelegalai build legal agents that summarize depositions, conduct research across case law, and draft work product. Patterns proactively finds the issues affecting only 1% of agent outputs, and helps resolve these silent regressions. Try Patterns → braintrustdata.link/agent-ob…
1
2
1,718
GPT-6 Astra is now available in Loop. Use it for pattern detection, eval creation, trace debugging, and other production investigations. Loop is built on Codex, so you get the same model capabilities on top of all your Braintrust context.
2
9
653
The Braintrust MCP server now exposes write tools. Your coding agent can author prompts, scorers, and classifiers, configure the Topics pipeline, build monitor views, create alerts, and run evals, all without leaving your agent environment. Read more → braintrustdata.link/write-to…
3
400
Your coding agent already knows how you work. It's set up with your repo, tools, and skills. Asking you to move part of that workflow into our in-product agent, Loop, is a high bar, so we have to give you something meaningfully better in return. To keep ourselves honest, we ran both options through a sample of 5 production investigations. Loop was 39% faster, but the insight quality was nearly identical. You should make your decision based on where you want to work next. Use Loop when you want fast, cited (down to the span) production investigations, and you're already in the UI. Especially useful for PMs and subject matter experts. Use the MCP + your coding agent when you want to take your insights into your repo to make changes. Read the full report → braintrustdata.link/loop-vs-…
2
7
798