BMW. BMW. BMW.
Replying to @samhogan
This is even funnier since we all bleached our hair
1
3
367
Mike Pollard retweeted
we're helping a customer spending $60k/mo move from OpenAI & Anthropic to open source models they use almost every model offered by the labs, so we needed to find replacements for all of them after generating evals, this is what we landed on new cost: $12k/mo, 80% savings
152
97
1,560
225,071
Clean up time :)
fable coming back to my codebase after 19 days of opus
2
98
Not a bad spot to work from
3
11
331
This is the easiest way to be validate that switching to an open source model will work for your flow without any interruptions.
Want to try GLM 5.2 in production but worried how it might change your product? Don’t worry, we got you: 1. Install Inference Gateway (docs.inference.net) 2. Keep sending traffic to your current provider 3. Gateway automatically starts sorting through your live data using an RLM to generate evals for your app. This takes ~24 hours. 4. Gateway starts mirroring live traffic to GLM 5.2 to run evals. Traffic is only mirrored - you’re still using your old provider in prod. 5. Once evals look healthy, you get a Slack notification letting you know it’s safe to switch. 6. Switch model identifier in your code to “glm-5.2” Congrats, you just saved 90% on your monthly token bill, and you own your LLM stack end to end.
3
179
Mike Pollard retweeted
META BRAIN MODEL | NEW AGENT BILL | FABLE BACK SOON? nitter.net/i/broadcasts/1nxeLLOWp…
7
6
67
49,250
We just launched agent signals! Now you can get notified when a user crashes out on your agent. Write any prompt and it will run on your trace data as it comes in, classifies it, and then optionally notifies you. Track your classifications over time to see trends in agent/user behavior. Check out this guide to get it set up: docs.inference.net/guides/me…
1
2
834
Mike Pollard retweeted
HALO is on the front page of HN :)
We’re releasing HALO Desktop 😇 it's the best way to find bugs in your agents import traces from Langfuse or Arize & have HALO create a report with failure modes give the report to your favorite coding agent to build it runs locally on your machine. 100% free and open source
2
2
23
5,044
Mike Pollard retweeted
We’re releasing HALO Desktop 😇 it's the best way to find bugs in your agents import traces from Langfuse or Arize & have HALO create a report with failure modes give the report to your favorite coding agent to build it runs locally on your machine. 100% free and open source
9
6
79
18,431
HALO just just got a lot more powerful. After you've integrated tracing you can now link your repo so the agent can recommend direct fixes based on all your trace data and the actual code itself. Right now it's free to try and a pretty easy way to improve your agent.
3
9
2,110
Put together a video of what we've been building over at @inference_net Check out how you can optimize your production agents with HALO and Catalyst
2
2
13
3,210
The keyboard from the @cursor_ai conference is pretty slick. Talks were great too
62
This is what I wanted agent observability to feel like. Traces that end in a code change.
1
1
5
1,226
It was great working with @oliveholistic to improve their product experience! If you want faster and cheaper inference at frontier model quality, shoot me a DM.
Specialized models are becoming a practical path to better AI UX. Olive moved from a frontier model to a custom model trained with Inference Catalyst for their food verdict workflow. After a user scans a product, the model now delivers near-instant verdicts on what to watch out for, making the in-store experience faster and more seamless while cutting inference cost significantly. Results: - p50 latency: 2,721ms → 591ms - p99 latency: 6,414ms → 998ms - time to first word: ~0.25s - inference cost: ~70% lower Great working with @oliveholistic on this! Full case study here: inference.net/case-study/oli…
2
60
Mike Pollard retweeted
Shout out to @samhogan @AmarSVS @francescodvirga @atbeme @mikepollard_dev and the rest of the Inference dot net team
A glimpse into our work with @inference_net thanks to @nvidia!
1
3
18
1,031
Install our SDK, collect traces, improve your agents.
3 weeks ago we open-sourced HALO this led to talking with dozens of teams running agents at scale we realized the current agent monitoring tools aren't built for the future that we so clearly see ahead of us today we’re releasing native OpenTelemetry-compatible agent tracing on @inference_net, powered by the same open-source core behind HALO
50
Mike Pollard retweeted
We're releasing Schematron V2, a family of Specialized Language Models for converting messy HTML to structured JSON frontier performance at 1/10th the cost Schematron V2 was designed in partnership with some of the largest web-scraping companies in the world to meet the demands of their heaviest workloads Schematron-V2-Turbo and Schematron-V2-Small are available today on @inference_net Get started: docs.inference.net/workhorse…
I found out today that two of the largest web scraping companies in the world are using a custom Llama 3 model we released last year to process millions of webpages per day. Schematron-3b: HTML -> JSON parsing Frontier quality at dirt-cheap prices. huggingface.co/inference-net…
6
8
74
13,753
Come train a custom model
Introducing Catalyst: a developer platform to monitor, train & deploy self-improving AI models built for teams operating AI products at scale Catalyst can automatically: - collect traces from your agents - curate training data & evals - train & deploy models on par w/ Opus 4.6
45