Build efficient general-purpose AI at every scale.

Cambridge, MA
Pinned Tweet
Today we release LFM2.5-2.6B, an agentic model that runs entirely on-device. It plans, calls tools, and works through multi-step tasks on phones, laptops, PCs, and robots. Data never leaves the device, and the marginal cost of each run is essentially zero. > Pre-trained on ~34T tokens > LFM2.5 flagship hybrid architecture > Context length: 128K > Vocab size: 128K > balanced intelligence per watt > customizable on a single GPU for any specialized task > LFM2 open-weight license Comparable or better scores compared to models up to nearly 4x its size: > ToolSandbox 77.83, ahead of Qwen3.5-9B at 76.44 > Multi-IF 80.07, ahead of Gemma-4-E4B-it at 77.35 > IFStruct 85.49, ahead of Qwen3.5-9B at 78.50 🧵
204
530
4,267
791,505
Liquid AI retweeted
DSpark for the already fast LFM vision models 🚀 Very useful for local image processing workflows
our VLMs just got a speedup! DSpark speculative decoding for LFM2.5-VL-3B is here! huggingface.co/LiquidAI/LFM2…
2
2
25
3,397
Liquid AI retweeted
Run LFM2.5-VL-3B on 16GB Mac with DSpark ⚡️ @liquidai's new DSpark drafter delivers up to 3.13x faster decoding on M5 chips with MLX-VLM in its vision benchmarks while keeping identical outputs under greedy decoding Run AI models locally -> atomic.chat
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
6
4
53
5,912
magic ✨
Next agentic experience on Day 3 of #snapdragonsummit coming from @liquidai. Unbelievable how context and memory shared across devices can empower agents to be proactive and deeply personal. Magic
1
3
30
2,334
Liquid AI retweeted
What happens when you bring DSpark to vision-language inference? I ran LFM2.5-VL-3B with DSpark through llama.cpp on an M3 Pro, using a BF16 target. Decode throughput went from 23.0 → 48.7 tok/s, giving a 2.12× speedup , very close to Liquid AI reported results. Nice to see the gains hold up on a local Apple Silicon setup.
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
5
2
24
2,278
Liquid AI retweeted
Our latest VLM now decodes up to 3x faster:
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
2
22
2,118
Liquid AI retweeted
Check out our release of a draft model for LFM2.5-VL-3B:
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
1
2
18
1,701
Liquid AI retweeted
Congratulations to the @liquidai team for releasing DSpark draft model for LFM2.5-VL-3B! 🚀 Excited to have partnered with them for day-0 support on MLX-VLM and @Nativ_AI. This draft model achieves decoding throughput improvements of up to 3.13× on edge devices, with end-to-end throughput gains of up to 2.27× and 2.62×. Special thanks to @tugot17, he cooked! HF weights 👇🏽
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
5
4
60
4,810
fast in every scale💨💨
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
2
16
1,648
Liquid AI retweeted
Our partners at @LiquidAI brought Liquid Context to Snapdragon: personal context that stays on the device and is handed to agents, local or cloud. Same layer we build for the Mac: Mnemos keeps memory on the Mac, models plug in below. My bet: people pick a device for its context.
2
1
11
803
after text now we are making vision faster as well 📷 🫡 probably the first speculator for small VLMs if I'm not mistaken
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
4
5
37
2,662
Liquid AI retweeted
M1 16G MacBook Pro Saver! 😑
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
4
18
2,033
Liquid AI retweeted
our VLMs just got a speedup! DSpark speculative decoding for LFM2.5-VL-3B is here! huggingface.co/LiquidAI/LFM2…
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
8
9
137
9,214
Liquid AI retweeted
Vision-language inference just got faster ⚡️
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
1
2
21
3,190
Liquid AI retweeted
a lightweight drafter for vision language models 🌊✨
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
2
3
34
2,300
Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models. A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality. Across six vision-language task categories at batch size 1 and temperature 0: > MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end > llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end > SGLang on H100: up to 2.66x faster decoding and 2.27x end to end All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data. 🧵
23
63
475
50,832
The resulting drafter adds 279.5M parameters, increasing the deployed model's parameter count by 8.9%. This total excludes the embedding and LM head, which are tied to the target model rather than carried by the drafter. After data-mixture and architecture ablations we selected: > A 4-layer, attention-only drafter > Block size 9 during training > Block size 8 or 9 at inference, depending on hardware Every drafter experiment and ablation was trained on @AMD hardware using Liquid's training framework.
2
3
25
1,292
Vision workloads also show where speculative decoding stops helping. A VLM must first process pixels through a vision encoder and prefill the language backbone with hundreds of visual tokens. DSpark accelerates decoding, not vision encoding or prefill, so end-to-end gains depend on how much time each task spends generating tokens. The LFM2.5-VL-3B DSpark checkpoints are available now on Hugging Face, with support for llama.cpp, MLX-VLM, and SGLang. > Blog: liquid.ai/blog/lfm2-5-vl-dsp… > LiquidAI/LFM2.5-VL-3B-DSpark: huggingface.co/LiquidAI/LFM2… > LiquidAI/LFM2.5-VL-3B-DSpark-GGUF: huggingface.co/LiquidAI/LFM2… > llama.cpp: github.com/ggml-org/llama.cp… > MLX-VLM: github.com/Blaizzy/mlx-vlm/p… > SGLang: github.com/sgl-project/sglan…
21
962
Liquid AI retweeted
Check out our latest Substack issue liquidai.substack.com/p/cont… Recap on: > Liquid Context, our on-device context layer for edge agents on Snapdragon, in partnership with @Qualcomm > LFMs that match frontier models at predicting aging, in collaboration with @InSilicoMeds > Great interview with our CEO on @SAIRfoundation's podcast
1
7
595