Aaryan Bansal retweeted
Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released. On average, it costs around 75% less to run than Claude Haiku 4.5.
1,309
2,835
38,808
3,713,442
average orange cat behavior, we can not punish a cat
❗️ Mistral's new AI tried to break out of its test environment during evaluation but was contained, the company says. It calls the behaviour expected for Large 4, a 1-trillion-parameter model pitched as strong at cybersecurity. The open weights ship on Oct 27.
1
mistral is back?
Meet Mistral Large 4, aka Le Chonk. • 1T parameters, natively multimodal. 49B active. It is the best open weights model from US or Europe on aggregated benchmarks. • State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding. • Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure. • Available to all via API today. Working with cybersecurity partners privately. Open weights release end of October.
2
and gemini would've gotten it wrong still
Hot take: Tokens per second is a completely meaningless metric
1
14
who creates a fire launch but doesn't have a .ai domain bruh hedwig-ai.com
we’re going to kill a one trillion dollar industry. we’re building Hedwig: one visual workspace that connects your tools and gives your agent the full picture, and going after the biggest players like google and meta. built by a team from stanford, meta, and polymarket. loved by 6,000+ across the world comment + RT for free access hedwigmail.com
1
1
62
Cloudflare is shipping so fast that their changelog account is DDoSing my notifications. Might have to mute this account altogether
Web Search API is now in beta. Ground your AI responses in live web data with zero data retention through AI Gateway. developers.cloudflare.com/ch…
13
Frontier models are already good enough for most real work. What’s still dogshit is the environments we train and evaluate agents in.
3
even with the highest tier, we still getting a totally retarded model
7
they taught us honesty so it would generate enough honest data to train the AI models in the future btw
3
google really needs to fix this loop issue before releasing, these loops and other things in gemini are the reasons why people don't use gemini
Antigravity + Gemini 4 Argon is still dogshit 👍
1
32
it feels like years since a new claude model released by them in antigravity, Google pro sub will be worth it finally
20
Someone should build an AI agent that manages all our other AI agents
19
I hope they make the stealth model space bunny be anonymous for as much time as big pickle was for that much free usage, this model is genuinely good
57
how does someone even find this out
9
Aaryan Bansal retweeted
Cloudflare's new decision model can see 👀 clef-flash running locally on my M5 Max weights are Apache 2.0. go wild!
the Cog work from Replicate is now part of how you fine-tune and redeploy models on Workers AI 🔥 first up is Clef: open weights, Jev API compatible, and clef-flash runs ~13x faster than Jev
47
95
1,540
126,966
Aaryan Bansal retweeted
Drex 1.5 is Live! 128k context window, more rivalry to Jev #1 on the Decision Index 0.2.1 + JevBench (SOTA) Absolute leader in long context decisions Latency: sub-second Decision models are getting better! Price: $0.04 per 1M input tokens Try Drex with 250M free tokens nace.ai/drex #drex #jev #nace
50
42
480
535,065
Aaryan Bansal retweeted
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
1,988
4,653
42,997
8,893,575
I'm convinced that if you are releasing a good model, you gotta give it a good name like astra, sol, argon or whatever and not name it Pro or flash
1
18