Most of what we call "agentic AI" is a string of tiny decisions. Which tool gets called. Which document matters. Which request gets routed where. Whether a message trips a guardrail. And we keep handing those decisions to giant models that write a paragraph of reasoning before picking option B. That's slow, it's expensive, and at high volume it adds up fast. I've built enough integrations to know the bottleneck usually isn't intelligence. It's latency and cost on the hundreds of small calls nobody budgets for. So I've been looking at Drex, a new small model from Nace.AI built for exactly this job. It doesn't generate text. It looks at the options you give it and returns calibrated probabilities over them in a single pass. No long reasoning trace, no token bill for thinking out loud. The numbers they're reporting: #1 on the Decision Index 0.1, beating Jev 1.13.0. 136ms response time. Up to 32K context. Up to 1.5x faster than Jev. Under the hood it's a diffusion architecture trained with reinforcement learning from adjusted feedback loops, which is a very different bet than scaling up another chat model. Where I'd actually use it: agent routing and tool selection, reranking search results, document classification, compliance and policy checks, security guardrails, and risk flags. Basically any step in a pipeline where you need a fast, confident pick instead of an essay. The fun part is they dropped it into Chess, DOOM, StarCraft, and Lemmings to show it making calls in real time. Worth a look if you want to see what 136ms decisions feel like. They're giving 250M free tokens to the first 10,000 builders. If you run agents or routing layers, test it on your own workload and see where it beats what you've got. nace.ai/drex #ad #AI #AIagents #LLM

Sep 25, 2026 · 8:12 PM UTC

12
5
11
14,339
Sort replies: Relevant Recent Liked
Replying to @RodmanAi
This is a really interesting approach to reducing the hidden costs in agentic systems.
1
129
Replying to @RodmanAi
32K context with that kind of response time is definitely worth experimenting with.
79
Replying to @RodmanAi
Curious to see how it performs on real production workloads
26
Replying to @RodmanAi
i started logging every tool call. the reasoning paragraph reads fine, but the tool choice is where it actually breaks.
28
Replying to @RodmanAi
i guess nobody wants to be the guy who routed the refund request to the tiny model :)
29
Replying to @RodmanAi
The focus on small, fast decisions makes a lot of sense for high-volume agent workflows.
20
Replying to @RodmanAi
Interesting architecture choice.
63
Replying to @RodmanAi
Tool selection seems like one of the most obvious places where this could help.
48
Replying to @RodmanAi
The governance version is funnier: every tiny decision is individually "allowed" and the chain is where nobody signed off. Small models for routing, sure. Just check the tool call against what the session was actually granted. A paragraph of reasoning is not an audit log.
3