Fast ML inference. Run top AI models using a simple API.

Palo Alto
DeepInfra retweeted
I used my entire usage limit for BOTH Codex and Claude Code (2 accounts) to tackle a local build, and the problem was never solved even though I used Fable 5.1 for intensive planning, parallel-agent supervision, and evals. I then went to @DeepInfra and got an API key for DeepSeek V4.1-Flash (US-hosted), and it is now CRUSHING this massive task Both Fable 5.1 & Astra could not. 50% of the way through & I only spent $1.15 so far. This is why all of a sudden the AI Cartel is sounding the alarm about AI safety. It's not about safety, it's about caping the competition. The numbers speak louder than perceived good will of the Effective Altruism mafia.
12
2
8
726
Encoder-decoder is back 😈!!! In DeepSeek V4.1 Flash the first 20 layers build the global KV that the next 20 read from, so prefill costs about half. Rolled out over the last 24h: throughput doubled, and we are approaching 1T tokens/day on OpenRouter. 30% off to celebrate. Enjoy. Cheapest on the market, as always.
1
3
34
1,311
Persimmon is a genuinely different idea: a model of how people actually talk, not another assistant. Proud to support @humansand on this launch with DeepCluster, a dedicated NVIDIA Blackwell cluster we deploy and operate. Excited to see where it goes. deepinfra.com/deepcluster
3
7
49
7,038
DeepInfra retweeted
For AI to work with us, it needs to understand us Today, we're introducing Persimmon, the first large-scale model designed to realistically simulate how people talk and interact
87
142
1,394
551,544
congrats to @humansand team on Persimmon release 💫
For AI to work with us, it needs to understand us Today, we're introducing Persimmon, the first large-scale model designed to realistically simulate how people talk and interact
3
3
23
1,872
Two Ling 3.0 flash variants are now live on DeepInfra, both day 0 with @AntLingAGI → Ling-3.0-flash-Fin — finance-tuned, 256K context. Source-grounded research across filings and earnings, valuation modeling, spreadsheet-aware output. → Ling-3.0-flash-VL — native multimodal, 1M context. Text, image and video, scoring 42 on the AA Intelligence Index. $0.06 in / $0.18 out / $0.012 cached
3
2
21
2,684
Ling-3.0-flash-VL is live on DeepInfra — day 0 with @AntLingAGI. deepinfra.com/inclusionAI/Li…
Today, we're open-sourcing Ling-3.0-flash-VL in BF16 and FP8. FP4 and INT4 are coming soon. Beyond visual recognition, it follows visual cues to: - Understand images, video, docs & UIs - Reason, search & verify - Use tools, check results & deliver
3
7
2,417
46% fewer errors 👀🚀
We got 46% fewer errors than the single best LLM across the 16 most used benchmarks (TerminalBench, LiveCodeBench, etc). Here's how that's possible and what each model can achieve when used optimally (every benchmarks misses the majority of model capabilities) 👇 Interactive Site: aifrontier.withmartian.com/ Academic Paper: arxiv.org/abs/2606.26836
9
828
Congrats to @AntLingAGI on the launch of their new model!
We’re open-sourcing Ling-3.0-flash-Fin, a finance-enhanced model for real-world workflows, and FinFIRST, an expert-built benchmark for financial search agents. Two open releases, one goal: making financial AI more accessible and verifiable.
1
1
12
1,685
The @Zai_org team is on a roll. Another open-weight drop, and a seriously good one — congrats to everyone who shipped it. Live on DeepInfra now -> deepinfra.com/zai-org/GLM-5.…
GLM-5.3 is now open-weight. Our most capable model for agentic coding and cyber defense is now available to download, run, and customize. Weights: huggingface.co/zai-org/GLM-5… Tech blog: z.ai/blog/glm-5.3
2
5
28
1,769
Day 0 support for GLM-5.3 on DeepInfra. 🚀 @Zai_org's latest flagship model brings major advances in complex coding, long-horizon agents, and cyber defense. GLM-5.3 is running in the US on @nvidia Blackwell GPUs with ZDR. $1.40 input · $4.40 output · $0.26 cached / 1M tokens
2
2
17
994
Video and audio, from one model. Wan3.0-Video from @Alibaba_Wan is now on DeepInfra: 30-second clips at 1080P, omni-modal reference (images, files, even web pages), and characters that stay the same character. $0.20/sec → deepinfra.com/Wan-AI/Wan3.0-…
2
4
9
19,333
congrats to @Zai_org team, big milestone!
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
1
1
21
1,642
Day 0 support for GLM-5.3-Flash on DeepInfra. $0.15 in / $0.50 out per 1M. $0.03 cached. @Zai_org's 320B-A18B multimodal model — 1M-token context, built for coding and long-horizon agents. Running in US on NVIDIA Blackwell.
5
6
41
2,058
Hybrid sparse + linear attention holds long-context accuracy without the usual compute curve on a 1M-token window. Chat, reasoning, tool calling, and JSON / structured output — all supported from launch. OpenAI-compatible endpoint. deepinfra.com/zai-org/GLM-5.…
1
3
442
DeepInfra retweeted
THE RISE OF THE DATA CENTER Here's how much revenue Nvidia $NVDA has brought in each quarter over the last couple of years from its Data Center business Q4 2019: $1B Q1 2020: $1.1B Q2 2020: $1.8B Q3 2020: $1.9B Q4 2020: $1.9B Q1 2021: $2B Q2 2021: $2.4B Q3 2021: $2.9B Q4 2021: $3.3B Q1 2022: $3.8B Q2 2022: $3.8B Q3 2022: $3.8B Q4 2022: $3.6B Q1 2023: $4.3B Q2 2023: $10.3B Q3 2023: $14.5B Q4 2023: $18.4B Q1 2024: $22.6B Q2 2024: $26.3B Q3 2024: $30.8B Q4 2024: $35.6B Q1 2025: $39.1B Q2 2025: $41.1B Q3 2025: $51.2B Q4 2025: $62.3B Q1 2026: $75.2B
37
57
518
55,972