OSS platform for AI agent evaluations, simulations & observability bhttps://github.com/langwatch/langwatch ➡️ Scenario Agent Simulations

Amsterdam
LangWatch retweeted
With LangWatch unique full SQL support, you can leverage the whole power of our platform structured data on your agent traces, cost analysis, tool calls details, sessions aggregations, events and much more, to cut and dice however you want And now `eval()` is a simply a function inside that SQL, adding intelligence to filter and score every single agent trajectory, for you to to build that perfect dataset You can export the matches as JSONL and use them as a test set, a post‑training set, or to distill a smaller model. It has never been easier!
1
1
1
132
LangWatch retweeted
Website: langwatch.ai/instant-evals Thanks to Jev, evals have never been faster and cheaper, while also above frontier models in human agreement, it's truly a breakthrough You can use it to find all sessions where users where frustrated, or where a new tool call would have saved tokens, or what top questions your users ask You can also use it to improve your own coding harness, for example processing all your past Claude Code sessions traces and finding all the times where you got annoyed at Claude and why Finally, as open models as now more powerful than ever before, you can use this for building training datasets for distilling to smaller models Possibilities are endless
2
1
7
1,392
LangWatch retweeted
Introducing Instant Evals on LangWatch Run any eval you can think of, on all your production trace history. Cheap, fast, and at scale. Thanks to @typesafeai's Jev model. Your production data is gold, and now you can mine all of it as fast as your ideas.
1
1
10
502
LangWatch retweeted
I have been using @VictorTaelin's solution for memory (OptMem) for 5 weeks now, replacing Claude Code's native memory, and since here at LangWatch we track everything more than anyone, I ran a full analysis on it I compared OptMem with the 3 months I had before working only with Claude Code Native memory, and luckily, in many of those sessions, Claude randomly didn't activate optmem at all so I kinda had an A/B test done for me tl;dr: 1. is it better than native? yes, but not by a lot, and in different ways (table below) 2. is memory useful to agents at all? a bit, mostly in long sessions 3. where does it help the most? avoiding rediscoveries and "traps", that is, mistakes that waste time 4. does it save on token cost? meh, a bit 5. are you continuing with optmem? yes so if you don't know how optmem works, check the gif below and their repo, it has an interesting approach of always keeping the memory a fixed size this alone is a for me a huge advantage, first and foremost because it gives me such a clear mental model and clear understanding of how my memories are being organized, it's both very structured and simply human readable, not free-for-all random markdown files, nor byte-compressed vector dbs or whatever second, it limits the damage radius of memory rot. Rot is a fact of life for anything AI agents touches which a human is not supervising. It's like vibe-coding an app through pure bruteforce, you can totally feel the slop compounding and everything rotting. Same happens with your memories, and the wrong ones actually may degrade your harness over time now not necessarily optmem is a better solution against rot, in fact because of pushing the append-only log it might be worse as agents seem to by default not use the forget functionality a lot (I'm investigating why), or as claude puts it: "How correction actually happens, per store. Native rewrites in place. 1,719 of the 2,081 write events are rewrites of an existing file, and MEMORY.md alone was rewritten 601 times. When an agent trips over a wrong file it usually fixes it, as in the thresholds file it rewrote with re-measured numbers. What nobody does is sweep: 39 percent of path tokens went stale when ADR-076 moved the tree, 68 of 260 files fell out of the index, and transcripts show live ls failures on dead index entries. So native rot is silent staleness from the world moving, and it is only corrected when tripped over. Nothing is ever lost or contaminated, and nothing ever stops growing. OptMem cannot rewrite, so it appends. 281 lines hold 3 explicit correction lines, one of which points at the wrong id, and 5 contradicting pairs with no marker at all. Both halves stay in the log and the tree does not know which one is the correction. That is how the refuted note ended up in wake and the right one did not. So OptMem rot is contradiction plus compression, and nothing corrects it in place by design." still however, somehow optmem has accumulated less misleading memories than native there are a lot of low hanging fruits I identified from this which will make optmem better for me, main one being waking it up also on subagents and considering loading it right from a CLAUDE.md import so agents adhere to it more often I'm also having a claude cleaning up my memory to remove rot, and plan to keep doing this analysis every month or so to keep the house in order there are a lot of other solutions out there, but memory is hard, and I cant only trust anything if I test it for months like I'm doing here, so that's why I'm favoring simplicity, control, and my own understanding over anything, so I'm sticking to optmem for now
18
11
112
10,468
LangWatch retweeted
One article, two companies, zero lines of integration code We co-wrote a piece with @LangWatchAI : same question asked twice, the cache-miss turn is 3.1s with a full llm waterfall, the cache-hit turn is 742ms and the llm span just isn't there That's what building on opentelemetry gives you betterdb.com/blog/one-trace-…
2
2
47
LangWatch retweeted
Replying to @LangWatchAI
Great visibility upgrade for managing AI development costs
1
1
57
LangWatch retweeted
If you want to stop BigAI from sneakily increasing token cost, support us on product hunt
This week we shared the Claude Code usage Tracker in LangWatch. Since so many positive responses we decided to do a proper launch on Product Hunt. Find the link in the first comment to support our team and let us know what you think! In short: You know what Claude Code costs per seat. But do you know which sessions burned the budget, which model did the work, or whether your cache paid off? LangWatch now tracks every Claude Code session: tokens, costs, cache hits, bash commands, MCP tool calls, file edits, and full terminal replays. Works with Claude Code, Codex, Gemini CLI, and opencode. Free for individuals. Start with: npx langwatch claude
2
2
13
698
Replying to @LangWatchAI
Amazing LangWatch for Claude Code finally answers the real question: where your AI dev budget actually goes.
1
2
85
This week we shared the Claude Code usage Tracker in LangWatch. Since so many positive responses we decided to do a proper launch on Product Hunt. Find the link in the first comment to support our team and let us know what you think! In short: You know what Claude Code costs per seat. But do you know which sessions burned the budget, which model did the work, or whether your cache paid off? LangWatch now tracks every Claude Code session: tokens, costs, cache hits, bash commands, MCP tool calls, file edits, and full terminal replays. Works with Claude Code, Codex, Gemini CLI, and opencode. Free for individuals. Start with: npx langwatch claude
15
2
32
1,269
LangWatch retweeted
Replying to @forgebitz
this "great, thanks" would have costed $79.32
2
1
63
5,068
LangWatch retweeted
Claude Code very smart. Claude Code make code. Claude Code also make bill. Big bill. -But where token go? -Did cache work? -What tool agent use? -Why agent spend 20 minutes fighting one login command? -Why agent call Bash 47 times like Bash owe him money? Before: Me open invoice. Me stare. Me close invoice. Me pretend me never saw invoice. Now: `npx langwatch claude` Every Claude Code session appear in LangWatch: → actual cost → cache reads and writes → every tool call → where the tokens went → full terminal session replay → exact moment agent lost the plot Also works with OpenAI Codex, Gemini CLI, and OpenCode. Free for individual use. AI coding agents become normal. Watching what they do before they sacrifice the village budget to the token gods should be normal too. Me now have observability. Me finally sleep.
1
6
12
538
Every Claude Code session now shows up in LangWatch with the real cost attached, alongside Codex, Gemini CLI and OpenCode. Token spend climbs month over month and nobody can say where the waste is coming from. Underneath that sits the part nobody wants to model: today's prices are VC-subsidised, and what the same usage costs once that ends. npx langwatch claude You get → spend per person and per team → what the subscription plans actually cost you → cache reads and writes → every tool call → full terminal session replay Which means the invoice stops being a mystery. You can see where the tokens went, whether cache actually worked, why the agent spent 20 minutes fighting one login command, and why it called Bash 47 times like Bash owed it money. Free for engineers on their own sessions. Teams get the shared view on top, with fine-grained privacy controls. langwatch.ai/claude-code-usa…
1
6
190
LangWatch retweeted
@LangWatchAI launched Langy, an AI engineer that turns production traces into scenario tests, evals, PRs, and CI proof. Its launch demo tested a customer-support voice agent. That is where AI-authored fixes need a hard verification loop.
1
1
1
233
Launching Langy inside LangWatch - for AI teams building agents, much more power onto catching issues, let Langy find these issues and write tests to improve before hitting production. Don't have a Reachy Mini @huggingface robot at home? Try a similar experience locally ;)
2
168
LangWatch retweeted
Introducing Langy: your own AI engineer We're bringing a Claude Code experience to the people who really know what the agent is supposed to do: the domain experts. Agents are doing legal, health, and financial work now, and the engineers building them are not lawyers, doctors, or bankers. The experts are, and until now their contributions died in a backlog. With Langy they describe the change in plain language, and what lands is a pull request with tests, best practices included. To celebrate, we gave it a robot body, and Langy is now our official mascot. In the video, we plugged it into Langy over MCP and asked it to test our customer-support voice agent. It wrote the test plan and kicked off the simulations itself. And if you're an AI engineer, this is more power for you too: your coworkers can experiment on the agent and improve it without pulling you from the hard problems, and their contributions arrive fully tested and reviewed by you, or they don't merge. We'll be dropping videos and tutorials on the next days and weeks on what it can do
3
1
8
498
LangWatch retweeted
Replying to @SergioEstebanCE
@SergioEstebanCE enjoying life with his ai engineer best friend: langy wish this could be you? @LangWatchAI
1
1
26
LangWatch retweeted
Building #AI agents sounds exciting. Running them in production is where things get serious ⚙️ @_rchaves_, Co-Founder of @LangWatchAI, dives into agent simulation testing, evaluation and LLMOps in practice 👀 Real AI. No fluff
1
3
5
274
LangWatch retweeted
Had a great and interactive session with @LangWatchAI for Defining Agent Quality for AI Engineers & PMs
2
2
12
380
LangWatch retweeted
Day 4 of LangWatch Skills Week: Since end of last year, we are seeing more and more AI enablement teams consolidating various Agent Development Lifecycle tooling from different teams, from homegrown evals to basic db logging, they now need a single solution for various teams, as agent quality becomes top priority. Maybe you had Langfuse for tracing, and DeepEvals for some local evaluators, but the collaboration from domain experts and PMs are still not happening, as a dev you still need to solve everything yourself, and never have time to add proper evals or agent testing because they always get pushed to "the next sprint". So we made a video on how to shorten this consolidation time to essentially zero, thanks to Skills, your coding assistant can now organize all your agent development tools so you can have best practices implemented and a single, collaborative platform for all the AI teams
1
1
2
255