OpenSand is an enterprise AI infrastructure platform that provides access to leading AI models through one API, making AI faster, easier, and more affordable.

Silicon Valley
Pinned Tweet
OpenSand is officially live 🚀 One API. Multiple AI models. Smarter token usage. Discover OpenSand 👇 🔗 MPost: [mpost.io/opensand-debuts-uni…] 🔗 NextFin: [nextfin.ai/en/articles/opens…] 🌐 opensand.ai #LLM #AI #Token
3
94
OpenSand retweeted
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:mimo.xiaomi.com/mimo-v2-6
378
871
8,811
1,454,494
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol Grok 4.7 scores +2 points over Grok 4.6 on the Intelligence Index, with strong performance on agentic knowledge work tasks. We evaluated the new model at xhigh reasoning effort. Congratulations to @SpaceXAI and @ElonMusk on the release! Key takeaways: ➤ Grok 4.7 joins the frontier of agentic knowledge work: Grok 4.7 gains +111 Elo over Grok 4.6 (high) on AA-Briefcase, our private benchmark for long-horizon agentic knowledge work, scoring 1657 Elo and placing it alongside Claude Opus 5 and Claude Fable 5.1 at the frontier. On GDPval-AA, it scores 1695 Elo, +90 ahead of Grok 4.6 (high). ➤ A leap in coding agent performance: Grok 4.7 (xhigh) with Grok Build scores 56 on the Artificial Analysis Coding Agent Index, up +9 points from Grok 4.6 (xhigh). Among models in their native harnesses, Grok 4.7 + Grok Build now ranks 4th, behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. ➤ Incremental performance changes elsewhere: Outside of agentic knowledge work, Grok 4.7 broadly matches Grok 4.6 (high) on the other Intelligence Index tasks. It improves on Terminal-Bench 4.0 (+4.5 percentage points) and GDP.pdf (+3.0 p.p.), with regressions on AA-LCR (-3.7 p.p.) and AutomationBench-AA (-1.1 p.p.). ➤ High token use across tasks: Grok 4.7's gains come with higher token usage. Grok 4.7 (xhigh) uses approximately 81k output tokens per Intelligence Index task, compared with 36k for Grok 4.6 (high) and 27k for GPT-6 Astra (max) - 125% and 196% more, respectively. Other model details: ➤ Context window of 500k tokens, unchanged from Grok 4.6 ➤ Pricing of $2/$6 per 1M input/output tokens with cache hits discounted to $0.50 per 1M tokens, matching Grok 4.6 ➤ Configurable reasoning effort spans low to xhigh. Our evaluation uses xhigh.
161
225
1,988
21,706,690
OpenSand retweeted
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
10,657
16,388
87,854
76,533,735
OpenSand retweeted
This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast.
9,214
35,400
340,134
138,064,434
Fable 5.1: A Quiet but Telling Upgrade – Anthropic quietly rolled out Fable 5.1 & Mythos 5.1 on Sep 1. Same base, but what truly sets them apart is the safety guardrail configuration. 🧵Thread below 👇 #Claude #Fable #Anthropic #AI #LLM
4
2
107
🧵 Bottom line – Better benchmarks, cheaper cache, protected reasoning. Not revolutionary, but a clear signal before IPO. Safety or strategic moat?
1
14
🧵 Key move – Digital signatures on chain‑of‑thought. New API accounts can't tamper with context without breaking it – blocks unsafe distillation. Now an explicit API rule, not just detection.
9
🧵 Cost – Cache read drops from $1 to $0.25 per million tokens (-75%), cutting agent costs up to 45%. Stronger performance, lower price – a solid combo.
8
🧵 Performance – Fable 5.1 tops 8 benchmarks, with 52.6% on Terminal‑Bench‑Science (vs 24.7% Fable 5, 22.4% GPT‑5.6 Sol). Agent reasoning takes a big leap.
43
🌊The OpenSand promo MV is finally here. One API. All the leading models. Lower token costs. Zero vendor lock-in. 🎬 Made with Seedance 2.0. 👇Watch the music video now! 👇 #AI #Developers #LLM #OpenSand #seedance #AIGC
2
105
Every token counts. Every dollar does too. OpenSand unifies the best AI models behind one API—so you pay less per token, integrate once, and never get locked in. One integration. All the leaders. Lower cost. Full control. 🪩 opensand.ai #AI #Developers
2
46
OpenSand retweeted
Create ten videos in Gemini Omni for FREE until 11:59pm PT tonight. Try it on the web or in the app and share your creations in the replies.
131
122
1,529
191,749
OpenSand retweeted
Appreciate it! Now let's understand the world through the eyes of Qwen3.8. 🥳
Replying to @arena
Qwen3.8-Max ranks #2 in Vision Arena scoring 1,305. Second only to Claude Fable 5 (High) which has only a 13pt lead.
90
122
2,185
142,124
OpenSand retweeted
We're open-sourcing Mixture-of-Kittens (MoK), our MoE training megakernel for NVL72s. It fuses all Mixture-of-Experts communication and computation into a single, fully deterministic kernel, and runs up to 2.37x faster than the strongest public baselines.
230
470
6,113
588,498
Announcing the Artificial Analysis Endpoint Accuracy Index, measuring how much of an open weights model's accuracy each serverless API endpoint preserves. We are initiating coverage with GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 coming soon Providers trade off accuracy to optimize for speed and cost. They quantize weights, write custom kernels and tune their inference stacks, and sometimes they simply ship bugs. We are bringing the rigor of our Artificial Analysis Intelligence Index to measuring endpoints, so developers can pick providers on accuracy, not just price and speed We benchmark each serverless endpoint against our own self-hosted reference deployment of the official weights, where 100% represents matching the reference. An endpoint is at reference parity when its result falls within the 95% confidence interval of the reference. Coverage is live for GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 accuracy coverage launching soon Key elements of the Endpoint Accuracy Index: ➤ Three areas, equally weighted: tool calling (BFCL-500, 500 questions, 3 repeats), scientific reasoning (HLE-250, 250 questions, 10 repeats) and long context recall (AA-LCR-25, 25 questions, 10 repeats). Each subset separates endpoints on the serving choices that drive accuracy differences, with repeats sized for tight confidence intervals ➤ Reference deployment: we self-host the official weights at the lab's recommended precision, following the lab's serving recipe, and publish the complete commands for each reference ➤ Inference parameters: we run the model's highest supported reasoning mode and each endpoint's highest supported output length and context window ➤ Confidence intervals: the parity test accounts for uncertainty in both the endpoint's runs and the reference's runs ➤ Rotating coverage: models enter once sufficient number of providers serve them and exit when a newer version in the same family supersedes them. We benchmark new endpoints as providers launch them and refresh all listed endpoints periodically ➤ Point in time: each result carries the date it was measured, with multi-day benchmarks dated to their final day Key results for GLM-5.2 ➤ Output token limits restrict accuracy. Restrictive limits cut responses off before the model finishes reasoning, and the most restrictive endpoints score half the reference or less on HLE-250 Key results for gpt-oss-120b ➤ Tool call handling separates endpoints. Providers parse and format tool calls differently, and some endpoints score 22% on BFCL-500 against 37% for the reference ➤ Serving configuration changes what the model does at the same requested settings. Some endpoints produce far fewer reasoning tokens at the same configured level, and restricted context windows truncate long context tasks Key results for DeepSeek V4 Pro ➤ DeepSeek V4 Pro endpoints are more in line with the reference. Majority of the endpoints are at reference parity, and DeepSeek's own first-party endpoint scores slightly above the reference
81
86
1,075
157,340
OpenSand retweeted
Today, we’re launching Alpamayo 2 Super, our frontier open reasoning model for autonomous vehicles. Beyond seeing, Alpamayo understands and reasons through the complex world - thinks before it acts. It’s a powerful backbone for robotaxis, trucks, shuttles, delivery vans, tractors and the long tail of mobile robots—billions of autonomous machines someday. We’re releasing it for commercial use under OpenMDW-1.1 so teams can inspect it, fine-tune it and deploy it—open models advance safety and security. The next wave of AI is robotics—and it starts with autonomous vehicles. Great work, Alpamayo team! blogs.nvidia.com/blog/alpama…
1,103
1,894
15,676
2,131,497
OpenSand retweeted
Gemini Spark can now tap into @GoogleChrome’s auto browse feature to take care of complex errands for you online 🔎 With your permission, Gemini Spark can use your logged-in accounts to take care of tasks like scheduling viewings for apartments you've saved or researching flight options and starting the booking process.
155
247
2,406
682,032
OpenSand retweeted
GPT-Live can listen while it speaks. To make that feel natural at ChatGPT scale, we rebuilt the voice stack from client to model. This new architecture keeps audio flowing continuously, so deeper reasoning and tool use don't interrupt the conversation.
560
950
10,938
1,346,807
OpenSand retweeted
An internal version of our next major model produced 10 new results on long-standing open problems in mathematics and theoretical computer science, using roughly $2,000 worth of tokens at GPT-5.6 Sol API rates.
580
986
13,907
1,812,891
OpenSand retweeted
Accurate 😂
3,800
4,929
47,812
13,654,749