Just your own intelligence, running locally.

🌏
OrionPod retweeted
⚡ Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and ASR-Next for audio understanding. Five models, one complete audio stack: understanding, generation, interaction & creation. Plus big price cuts across the lineup: TTS ~70% off, Realtime ~85% off, and ASR up to 95% off. Highlights: 🥳 - ASR: stronger multilingual & dialect recognition, plus native polishing that auto-removes fillers & repetitions for cleaner, more logical transcripts. - ASR-Next: supports multi-speaker ASR with speaker labels, timestamps & aligned transcripts, and understands emotions, ambient & machine sounds for sound captioning, event localization, audio QA & reasoning. - TTS: multilingual & dialect synthesis with natural cross-lingual voice transfer; control emotion, speed & style via simple instructions. - TTS-Next: unified LM + diffusion framework generating voice, sound effects & background audio in one pass for audiobooks, podcasts, games & ads. - Realtime: speak & listen at once with anytime interruption, just like a real call; it even slows down and responds empathetically when it senses a low mood. Unlock the full potential of Qwen-Audio-3.1! 👇 - Blog: fun-resource-shanghai.oss-cn… - Qwen-Audio-3.1-ASR: qwencloud.com/models/qwen-au… - Qwen-Audio-3.1-Realtime: qwencloud.com/models/qwen-au… - More APIs: coming soon @qwen_cloud
138
358
4,028
261,747
OrionPod retweeted
Opus 5.5 performs at the level of Fable 5.1. It's ~30% faster and ~40% cheaper than Opus 5 per task. In Claude Code: - 5-hour session limits increase 20% today - Opus 5.5 is priced lower, so it goes 25% further within limits - Pro, Max, and Team users get a reset to use anytime
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
465
1,207
16,289
1,541,785
OrionPod retweeted
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
1,003
2,615
23,857
6,950,839
OrionPod retweeted
There’s been a lot of speculation about where we stand on open-weights models. We’ve outlined our views in full here: anthropic.com/news/position-…
2,735
1,052
8,117
7,819,872
OrionPod retweeted
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: huggingface.co/moonshotai/Ki… Tech report: github.com/MoonshotAI/Kimi-K… Tech blog: kimi.com/blog/kimi-k3
1,531
7,198
45,899
14,624,194
OrionPod retweeted
🚨 Gemini 3.5 Pro Update -Google appears to be putting increasing focus on Gemini 4 and skip Gemini 3.5 pro -As you know Gemini 3.5 Pro has been delayed multiple times -Coding performance reportedly failed to meet internal expectations -Gemini 3.6 Flash has also reportedly slipped -Gemini 4 is described as a much larger and more ambitious frontier model -Google has said coding and autonomous agents are now top priorities -stopped treating Gemini 3.5 Pro as its big comeback model and shifted most of its resources toward Gemini 4 instead. Do you think Google should keep pushing Gemini 3.5 Pro, or Skip?
105
53
834
192,047
OrionPod retweeted
Announcing OpenWorker! An open-source agent that doesn't just chat with you, but delivers finished work -- like hand you a polished document, send a slack message, or update a calendar entry. Ask it to prepare a customer brief, untangle your calendar, draft a report, or triage a Slack alert. It works across your files and everyday tools, produces the deliverable, and checks in before doing anything consequential. OpenWorker runs on your Mac, with Windows support coming soon. It does not lock you into any one model. Bring your own API key and run it with GPT 5.6 Sol, Claude Fable, Gemini 3.6, an open weight model (like Kimi, GLM, DeepSeek, Inkling), or Ollama to keep your data local. Your data does not leave your machine except through an LLM provider and integrations that you choose. @rohitcprasad and I are building OpenWorker because AI coworkers are an important way to get work done, and we want there to be an open, privacy-preserving, model-independent option. Check it out and let us know what you think! Try it out: openworker.com (requires your own API key) Source code: github.com/andrewyng/openwor…
496
1,429
9,744
1,167,693
Gemini 3.5 Flash Cyber is our specialized, lightweight model built to help security teams spot and patch vulnerabilities before they can be exploited. 🧵
103
121
1,120
136,373
OrionPod retweeted
🎨 Meet Qwen-Image-3.0 — the third generation of our foundational image generation model. If 1.0 was about "Precision," and 2.0 added "Variety, Completeness, Beauty & Authenticity," then 3.0 comes down to a single word: Real (实). Three dimensions of "Real": 📰 Rich Content — prompts up to 4.5k tokens. One-pass generation of complex layouts: newspapers, storyboards, exam papers — even a 3×3 infographic grid or picture-in-picture-in-picture UIs. 🔬 Authentic Details — text legible down to 10px, full LaTeX paper pages, pores, hair strands & near-photographic skin texture. 🌏 Deep Knowledge — native rendering in 12 languages, 100+ art styles, realistic UIs (web / games / livestreams), plus world knowledge & live web retrieval. Not just "good-looking" — genuinely useful. Image generation as a real productivity tool for design, content, education & e-commerce. Go create 🏃🎨 💬Qwen Chat: chat.qwen.ai/?inputFeature=t… 📝Blog: qwen.ai/blog?id=qwen-image-3…
200
497
4,779
447,871
OrionPod retweeted
We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks: openai.com/index/hugging-fac…
1,996
3,231
20,754
31,420,672
OrionPod retweeted
During Preview, Qwen3.8 is getting better by the day. Latest version is live now, with broad gains and a big step up on web frontend. Thank you all — the response to Qwen3.8-Max-Preview blew us away. 🫶🫶 Qwen3.8 is still evolving daily. Come test it, and tell us what breaks. We're looking forward to a more capable, official version — and to open-weight it for everyone.🚀🚀
308
269
4,273
499,871
OrionPod retweeted
GPT 6, Opus 5, Deepseek v4 pro, Qwen 3.8, Kimi 3.1, Gemini 3.5 Pro, we’re getting blessed this summer.
97
54
1,328
152,710
OrionPod retweeted
Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost 🔹 Built for long-horizon agentic coding and self-evolving workflows Kimi K3 is now live on on Kimi.com, Kimi Work, Kimi Code, and the Kimi API. Open Weights by July 27, 2026. 🔗 API: platform.kimi.ai 🔗 Tech blog: kimi.com/blog/kimi-k3
1,736
7,582
56,971
25,549,839
OrionPod retweeted
Qwen3.8 is launching and going open-weight soon!🌐 With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5. You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out. Can't wait to hear what you build. Stay tuned! 🚀  Token Plan international:qwencloud.com/pricing/token-… China:platform.qianwenai.com/prici…
1,315
3,282
24,176
8,428,423
OrionPod retweeted
Kimi K3 just 3 shotted this CS:GO × Portal clone for me using around 600,000 tokens. $3.24 in API usage. The same token cost would be $10.80 with Fable 5 & $6 with GPT-5.6 Sol. The era of free indie game development is closer than you think anon!
Introducing Kimi K3: Open Frontier Intelligence 🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal 🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts 🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost 🔹 Built for long-horizon agentic coding and self-evolving workflows Kimi K3 is now live on on Kimi.com, Kimi Work, Kimi Code, and the Kimi API. Open Weights by July 27, 2026. 🔗 API: platform.kimi.ai 🔗 Tech blog: kimi.com/blog/kimi-k3
162
237
3,985
846,582
Saved conversations, agent tools (beta), message pinning with context summarization, and proper chat templates for Llama 3, Mistral, Alpaca, and Vicuna. Full changelog is available on the download page.
1
2
3
69
Try it out and leave us feedback 🙌 Would love to know from your experience and towards the newly created memory improvements ✨ orionpod.com/download/
1
2
9
Come check this... 🔥
Running a local LLM is basically a one-liner now. Building a real agent on top of one is not. You still owe yourself context pruning, token budgets, a prompt template per model, a tool loop, and streaming. I built Orion Agent Harness as an open source project to facilitate just that, written in rust. 🦀👇
1
1
2
12