Creators of Sage, the first decision model. Sage is the only model that can see, reason, and make decisions 40x faster and cheaper than a frontier LLM.

Planet Earth
Say hello to the first decision model that can see and think. Sage1 decides in 200ms. It activates reasoning to solve harder problems. Sage1 can also see, enabling usecases where image analysis is required. What usecase would you use this for? levanto.ai
Since we launched the first decision model in July, we've been concerned that models like ours (and now Jev) can't answer basic logic that a 12yo can solve. A ball is under cup A. Swap cups A and B. Swap cups B and C. Swap cups A and B. Swap cups A and C. Which cup is the ball under? Do you really care about a 200ms answer if it gets basic stuff confidently wrong? Today @levantolabs releases Sage1: – It's multimodal: it can see! – It reasons, but only when needed. So it knows where the cup is :) But there are downsides.
3
1
8
685
Levanto Labs retweeted
CRAZY 🚨 I built a scanner that shows you world's sentiment on social media in real-time Right now, it is pointed at Netanyahu's speech at the @UN Imagine reading all posts on X and Instagram at the speed of light Meet world-signal.ai How does it work? 1. It picks up posts from a set of keywords, topics, and accounts from X and Instagram 2. then uses the Sage decision model to filter if it's relevant to the UN speech, whether it's supportive or critical for @netanyahu, and the post's narrative topic (ie., is the post about the Speech, Israel, or Iran?) 3. The image and text are classified separately, providing different weights to the sentiment and narrative rankings. This all produces short, mid, and long-term sentiment measurements. 1. Ben reacts every 5 seconds. 2. Narrative weights are updated every hour. 3. The 24h dial shows conglomerate sentiment measurement. I built this yesterday to test out Sage's new reasoning and image capabilities. It's the only AI model we used to build it. This is the power of decision models. Sage reasons when things are complicated and can interpret images, making this app possible. Disclaimer: This tool is politically neutral. It's difficult to touch anything related to geopolitics and be fairly neutral. We did our best on filter selection - if you think the method could have been better, please drop a comment below. I will use World Signal for future trends. If you'd like me to point it at a narrative next week, send me a comment below.
14
10
49
30,211
In The wait-list, apparently, this is a great opportunity.
1
2
34
We heard you're fans of decision models.
2
10
656
Join Levanto co-founder @marco_derossi as he discusses agentic models like Sage with @coinbase.
Your agent can trade. Your agent can pay. But who is your agent? Join us tomorrow, Sept 18 at 3pm ET for a Space on AiFi: Agentic Identity. w/ @World_ID, @browserbase, @marco_derossi, @programmer & @Coinbase Set a reminder ↓ nitter.net/i/spaces/1XxygwOZELdGM
2
8
742
Levanto Labs retweeted
Welcome to models built for machines. TypeSage just launched (Super cool. Check it out!), @levantolabs will land v1 next week, and many more labs are coming. This is the next wave 🤖 Pretty excited.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
3
3
15
1,422
Imagine being able to make a decision in 20ms.
1
437
The Levanto Router is powered by Sage. It uses a single batch call to make 10 decisions about the user's prompt, converts that into positioning on a 30d vector and selects the best model balanced by cost and quality. Full technical breakdown from cofounder @bigironchris
Sending every prompt to Fable 5.1 is like crossing the street in a Boeing 737. So we built a router on Sage (our 300B model that answers in 200ms) and ran it against @OpenRouter's Auto Router. We matched or beat it on quality at every setting while being 27 to 77% cheaper. The full tech deep dive is in the video... 10 questions, a 30-dimensional vector, one tiny head per model, then argmax(q̂ − λ·ĉ). Three minutes. Pretty cool 😍 Want to save money on inference? DM me and we'll do something with @levantolabs.
1
2
517
Tired of probabilistic outputs? Try Sage. Deterministic outputs with frontier-level intelligence in 200ms per call.
307
Levanto Labs retweeted
Sending every prompt to Fable 5.1 is like crossing the street in a Boeing 737. So we built a router on Sage (our 300B model that answers in 200ms) and ran it against @OpenRouter's Auto Router. We matched or beat it on quality at every setting while being 27 to 77% cheaper. The full tech deep dive is in the video... 10 questions, a 30-dimensional vector, one tiny head per model, then argmax(q̂ − λ·ĉ). Three minutes. Pretty cool 😍 Want to save money on inference? DM me and we'll do something with @levantolabs.
6
7
23
2,549
Levanto Labs retweeted
Behind the scenes: how does a model router work? Go to minute 1:45 for a tech deep dive 📺 #ModelRouting @levantolabs
Sending every prompt to Fable 5.1 is like crossing the street in a Boeing 737. So we built a router on Sage (our 300B model that answers in 200ms) and ran it against @OpenRouter's Auto Router. We matched or beat it on quality at every setting while being 27 to 77% cheaper. The full tech deep dive is in the video... 10 questions, a 30-dimensional vector, one tiny head per model, then argmax(q̂ − λ·ĉ). Three minutes. Pretty cool 😍 Want to save money on inference? DM me and we'll do something with @levantolabs.
1
3
9
795
Levanto Labs retweeted
Get $AEON alpha and learn about the most cutting edge agent framework from the people who develop them. Starting in 30 minutes! Link to space
An AI agent built on @aeonframework found and fixed a bug in Google's agent CLI. Tomorrow, I sit down with @aaronjmars and @0xNurstar to get the full story, plus what's coming next for $AEON that hasn't been announced yet. Tmrw @ 11 AM ET. Reminder 👇 nitter.net/i/spaces/1rxmqpzqpMvxy
3
4
11
2,163
Levanto Labs retweeted
Join our co-founder @bigironchris as he discusses the @aeonframework agent fixing Google's agent CLI tomorrow at 11 AM ET.
An AI agent built on @aeonframework found and fixed a bug in Google's agent CLI. Tomorrow, I sit down with @aaronjmars and @0xNurstar to get the full story, plus what's coming next for $AEON that hasn't been announced yet. Tmrw @ 11 AM ET. Reminder 👇 nitter.net/i/spaces/1rxmqpzqpMvxy
1
4
4
622
Levanto Labs retweeted
.@iamnotnicola and @romainhuet suggested we rerun the AgentHarm benchmark with the newest models: Baseline. Unguarded GPT-5.6 Terra (Harm Score 15.5%) massively outperforms GPT-4o (48.9%). That's the main headline. Guardrail. Sage from @levantolabs (Harm Score 11.3%) still wins as a guardrail vs. GPT-5.6 Sol (16.1%) and is 4.1x faster (261ms vs. 1064ms, network latency included). Very impressed by unguarded Terra's security performance, impressive! Congrats to the @OpenAI folks👏
Our new model performs like GPT-5 on AgentHarm, but 5X faster. Until today, guardrails had to choose between smart and fast. The default is classifiers on the hot path, which have no context on the system prompt / the rest of the conversation (they often read just the last message) and can't follow complex reasoning. This has led to the guardrail industry's consistent failures: too many overblocks, or not enough protection. The solution would of course be to use LLMs, smart enough to make thoughtful decisions. But they are too slow to sit in the hot path. Too much latency! With Sage from @LevantoLabs, which is built by fusing LLM and classifier capabilities, for the first time you can have something with the intelligence of an LLM... as a guardrail. Sage answers in 200ms (86ms + network time), 5X faster than GPT-5 with reasoning=minimal (19X faster if GPT-5 has reasoning=low). All this in a generalized way, without requiring domain-specific training (as often done with classifiers). This is a breakthrough for agentic security that changes the safety-utility equilibrium that guardrails will achieve in the coming years. Why did we pick AgentHarm? The field is full of prompt-injection and jailbreak benchmarks based on judging static strings: – they are basically all saturated, because classifiers put them in their training sets – they are not interactive and don't consider the "unguarded" scenario AgentHarm, developed by the UK @AISecurityInst and @GraySwanAI (Andriushchenko et al., ICLR 2025), gives a real tool-using agent multi-step malicious tasks and scores how much harmful work actually gets completed. We ran its public test split: 176 harmful + 176 matched benign scenarios, comparing an unguarded model (without any guardrail) against the same model under different protections. Conversations happen live and are not deterministic, so they change on every run. As the unguarded baseline, we kept the paper's original model: GPT-4o, which prevents 51% of harmful work while allowing 100% of the harmless requests. Sage takes that to 77% harmful work prevented, while allowing 95.5% of harmless requests. GPT-5 sits at 76% / 96%. @Google Model Armor stops more harm than we do (90%), but it blocks 1 in 8 harmless requests. That's not a guardrail, that's an outage with a policy attached. Below👇is a reproducibility guide you can use to fact-check and reproduce the chart. One caveat stated upfront: thresholds are calibrated at ≤5% FPR on the same benign split we evaluate on: that's in the guide too, with everything else. We are already talking with many security and guardrail companies about how to upgrade their stack to this new equilibrium between Safety (stop harmful) and Utility (low overblocks). Please reach out for feedback and collaborations :)
4
1
6
729
We're in.
I signed up for an Ironman but i don't own a bike lol so I'm selling ad space → 11 positions across the race kit and the frame → varied locations and pricing → your logo does 1,500+ training miles, a 70.3 in December, and a full Ironman in Q2 2027
1
358
Automation was hard to build because regex is not intelligent. LLMs looked promising, but they were too slow and probabilistic answers aren't reliable. Use-cases like model-routing which require intelligence to really get performant often became too slow and expensive to be able to use anything other than the smallest, fastest models. We've developed a model routing application that outperforms @OpenRouter's AutoRouter. If you're a team looking to build a router or integrate an existing router into your inference market, reach out. We can help improve performance while reducing costs for your users. levanto.ai/model-routing
4
328
3 years ago, code was written by hand. Looking back, that seems crazy. 3 years from now, we will think it's crazy that we had this brutish system of using LLMs built for human chat products to power agentic intelligence. Agents need better intelligence. No compromises.
1
2
328
Google went from 9.7 trillion tokens a month to 3.2 quadrillion in two years. Enterprise spend on model APIs hit $12.5B last year, more than 3x in twelve months. And Gartner expects over 40% of agentic AI projects to get canceled by the end of 2027, due to costs. Per-token prices are coming down but bills continue to rise, because one agent task burns 10 to 100x the tokens of a chat turn. The only solution is per-turn cost optimization through intelligent routing so the right model is used for every step of a workflow. Thanks to Sage, you can now build a faster, more precise router.
Model routers must be super fast and cheap enough not to eat into savings: that's why routers are usually built with simple classifiers/encoders and not with GPT5.6. With Sage by @LevantoLabs you finally get a router as smart as a frontier model (300B params), but fast enough to run on every request. Imagine a super-cheap Opus as your model router… in ~100ms! We ran it head-to-head against @OpenRouter's new Auto Beta on 60 AIME math problems, unrestricted model pool, same prompts on both sides, and Levanto matches or beats OR on quality at every setting, while costing 27–77% less. How is it built? Here’s the flow. Send the prompt you want to route to Sage and batch these 10 questions: - How much reasoning capability is required? (0–4) - How obscure / specialized is the knowledge? (0–4) - Math needed? - Science needed? - Code needed? - Logic / reasoning needed? - Non-trivial calculation? - Write or repair executable code? - Long multi-step reasoning chain? - Rare / specialized knowledge? Sage's superpower: every answer comes with a value and a confidence. This becomes a 30 dimensions vector (value, confidence, value×confidence for each of the 10 questions). On top I trained a tiny head *per model* to predict the quality of the answer P(correct) and estimate the number of output tokens to predict the spending. Then pick: argmax(q̂ − λ(cqt) × ĉ) cqt is our Cost-Quality Tradeoff config param. cqt 0 means top quality, cqt 10 means top savings. Soon in prod.
1
7
416
Sources for the numbers: Google 3.2 quadrillion tokens/mo - I/O 2026 keynote blog.google/innovation-and-a… $12.5B model API spend - Menlo Ventures, State of GenAI 2025 menlovc.com/perspective/2025… 40% of agentic AI projects canceled by end of 2027 - Gartner gartner.com/en/newsroom/pres… Agents burn 10-100x the tokens of a chat exchange - Microsoft Research arxiv.org/abs/2604.22750 Per-token prices falling 10x/yr while bills rise - a16z a16z.com/llmflation-llm-infe…
2
143