Visual Intelligence for Developers

Santa Clara, CA
Pinned Tweet
Introducing VLM Run Gateway: run open-weight VLMs, OCR VLMs, embodied VLMs, and specialized ViTs in one OpenAI-compatible API. You don't need a trillion parameters to parse documents, understand images, or summarize videos. When every model runs on one API, changing models is as simple as changing a parameter. The visual frontier just moved. Your AI bill should too. $ uvx vlmrun gw models
5
3
10
64,725
VLM Run retweeted
Jev can't see yet, but it turns you can get @googlegemma to play Snake directly from pixels. NO this isn't yet-another Jev demo that uses game-state. It's seeing images, making decisions and responding e2e over websockets in <240ms p99 at ridiculously low costs (~$0.00007/image). We just added @typesafeai SDK compatibility to our @vlmrun gateway where you can try out a whole bunch of VLMs with System One support. Check it out, and it's free to try! Happy hacking on the weekend!
2
1
11
719
I made a computer vision tool for rock climbing analysis in 3D using iPhone LiDAR! Having learned a lot from sharing my previous rock climbing demos, I realized that a lot of rock climbing analysis is well-suited for 3D. Even something as simple as supporting videos where the person filming moves with the climber requires 3D information. To get the depth information, I used my iPhone 15 Pro’s LiDAR depth sensor through my local iPhone app. I found that the depth measurements help a lot. I think the holds activation is better, and I like the final view of all of the holds in 3D. It’s also interesting to see the distance traveled in meters. Plus, it looks cool and it feels like a video game 🙂 In short, I think this new demo is an improvement in that climbers can see the real-world distance traveled and a visualization of which hand and foot activated which hold. I recorded the video and depth measurements from my iPhone app, and I ran the rest of the analysis on my computer. I used ViTPose+ Large for pose estimation and SAM 3.1 to segment the holds, both models accessed through the @vlmrun Gateway. Let me know what you think! The analysis code is open-source on GitHub: github.com/jeremyipark/visio…
26
79
835
59,566
We ran @nvidia's blowout Q2 10-Q through glm-ocr on the @vlmrun gateway. 61 pages · 2,086 tok/s · 29.54s · $0.019 2K+ toks/s, and less than 2¢ to read a full quarterly filing.
1
2
7
694
One call, no pipeline, no per-page OCR bill. Try it today: $ uvx vlmrun gw chat <input>.pdf -m glm-ocr
2
173
Introducing VLM Run Gateway: run open-weight VLMs, OCR VLMs, embodied VLMs, and specialized ViTs in one OpenAI-compatible API. You don't need a trillion parameters to parse documents, understand images, or summarize videos. When every model runs on one API, changing models is as simple as changing a parameter. The visual frontier just moved. Your AI bill should too. $ uvx vlmrun gw models
5
3
10
64,725
Same test on video: MVBench score vs. cost per 1k clips. Three models land in the optimal quadrant (70+ score, under $10 per 1k clips), including two open-weight models on Gateway: - qwen3.8-27b (75.5% at ~$5.50) - muse-glimmer-30b (74% at ~$1.50) muse-glimmer-30b is the cheapest in the quadrant, delivering high accuracy at ~1/6 the price of Gemini 3.7 Flash.
1
4
115
We're at @foxglove #actuate26 to connect with fellow roboticists to learn more about about how teams are scaling robotics datasets, RL, and real-world deployments. If you're still hand-labeling > 50% of your robotic data, let's chat. We're releasing a new auto-labeling VLA/VLMs that you'd definitely be interested in trying out. DM @spillai or @dineshredy if you're around
I'm feeling incredibly excited about #Actuate26 this week, and grateful to our team who have dedicated the past few months to making sure the event is a success. It's hard to believe what this little conference has become! We're officially SOLD OUT - 1500 tickets! 3x last year. We've got 40,000 sqft of robots, 40 speakers, 2 stages, 2 happy hours, @UFBots, and 15+ community events happening during the week. See you all tomorrow!
2
3
445
Gemini 3.7 Flash is now available in our Orion 2 visual agent harness! Its improved reasoning and multi-step planning lead to stronger agentic performance within Orion 2 across visual workflows.
Gemini 3.7 Flash is here. It’s stronger for coding, knowledge work, and web development. 🧵
1
1
6
441
Qwen3.8-27B is now available in our Orion 2 visual agent harness! Orion 2 pairs Qwen3.8-27B’s native multimodal capabilities and improved agentic planning with an expanded visual toolkit, turning visual workflows into inspectable, executable programs.
We promised open weights for Qwen3.8. Now, time to meet them! 🎉 ⚡ Qwen3.8-27B: - A native multimodal dense model. With just 27B parameters, it outperforms Qwen3.7-Plus overall and shines in real-world coding & office workflows. - 262K native context, easily extendable to 1M tokens via YaRN. - Built for builders. Highly efficient, high-quality, and licensed under Apache 2.0. 🚀 The open weights for Qwen3.8-2.4T-A95B (Max-level) have also been released recently. Whether you're shipping lightweight applications with Qwen3.8-27B locally or building agents with Qwen3.8-2.4T-A95B, they're yours now! Download, deploy, and build something we haven't imagined yet. 👀👇 - Hugging Face: huggingface.co/collections/Q… - ModelScope: modelscope.cn/collections/Qw…
3
1
7
340
Try Orion 2 with Qwen3.8-27B for your visual workflows today: chat.vlm.run
1
73
VLM Run retweeted
Gemini Robotics ER 2 is now available in our Orion 2 visual agent harness! Orion 2 pairs Gemini Robotics ER 2's real-world understanding with an expanded visual toolkit, turning embodied reasoning into composable, inspectable programs.
One brain. For any robot. 🤖 We’re launching Gemini Robotics 2: our next-generation physical AI bringing full body intelligence to humanoids, advanced dexterity, multi-robot teamwork and more.
1
4
11
998
Gemini Robotics ER 2 is now available in our Orion 2 visual agent harness! Orion 2 pairs Gemini Robotics ER 2's real-world understanding with an expanded visual toolkit, turning embodied reasoning into composable, inspectable programs.
One brain. For any robot. 🤖 We’re launching Gemini Robotics 2: our next-generation physical AI bringing full body intelligence to humanoids, advanced dexterity, multi-robot teamwork and more.
1
4
11
998
Check out these example chat threads from our initial testing of Gemini Robotics ER 2: 1. Crop segment of adding rice to rice cooker + analyze frame: chat.vlm.run/c/40b5edfb-6d15… 2. Extract 16-keyframe grid + detect water bottle: chat.vlm.run/c/a8542517-0021… 3. Extract frame + segment lettuce + generate 3D reconstruction: chat.vlm.run/c/494c7c65-4aa1…
1
4
82
Kimi K3 is now available in our Orion 2 visual agent harness! Orion 2 pairs Kimi K3's multimodal understanding with an expanded visual toolkit, compiling multi-step tasks into deterministic, inspectable programs.
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: huggingface.co/moonshotai/Ki… Tech report: github.com/MoonshotAI/Kimi-K… Tech blog: kimi.com/blog/kimi-k3
1
2
8
1,319
Run Orion 2 with Kimi K3 for your visual workflows today: chat.vlm.run
2
58
Gemini 3.6 Flash is now available in our Orion 2 visual agent harness! Orion 2 augments Gemini 3.6 Flash's multimodal understanding and token efficiency with an expanded visual toolkit, composing multi-step tasks into interpretable, executable programs.
We’re rolling out three new models to make AI agents faster, smarter, and cheaper at scale: 🔵 Gemini 3.6 Flash: It uses fewer tokens than 3.5 Flash to deliver higher quality work at the exact same cost. 🔵 Gemini 3.5 Flash-Lite: A fast, cost-effective option for everyday tasks like processing documents and agentic search. 🔵 Gemini 3.5 Flash Cyber: A cybersecurity model built to find and patch critical software vulnerabilities.
1
1
4
1,157
Try Orion 2 with Gemini 3.6 Flash for your visual workflows today: chat.vlm.run/
1
72