100x | pumping MPI up co-founder @xai | @uoft

San Francisco, CA
Last day at xAI. xAI's mission is push humanity up the Kardashev tech tree. Grateful to have helped cofound at the start. And enormous thanks to @elonmusk for bringing us together on this incredible journey. So proud of what the xAI team has done and will continue to stay close as a friend of the team. Thank you all for the grind together. The people and camaraderie are the real treasures at this place. We are heading to an age of 100x productivity with the right tools. Recursive self improvement loops likely go live in the next 12mo. It’s time to recalibrate my gradient on the big picture. 2026 is gonna be insane and likely the busiest (and most consequential) year for the future of our species.
556
715
8,827
3,469,641
Energy is the common denominator of progress and innovation
15
11
319
110,121
accelerating towards a compute driven economy
13
7
305
77,228
Grok has the most intelligent voice
Today, we're excited to launch the Grok Voice Agent API, empowering developers to build voice agents that speak dozens of languages, call tools, and search realtime data. x.ai/news/grok-voice-agent-a…
20
14
319
60,463
we’re hiring. macro harder
Let’s see if @Grok 5 can beat the best human team @LeagueOfLegends in 2026 with these important constraints: 1. Can only look at the monitor with a camera, seeing no more than what a person with 20/20 vision would see. 2. Reaction latency and click rate no faster than human. Join @xAI if you are interested in solving this element of AGI. Note, Grok 5 is designed to be able to play any game just by reading the instructions and experimenting.
308
372
3,665
1,635,907
Very excited about the Agent Tool API and Grok 4.1 Fast! Developers now get the native Grok tools: code execution, web search, real-time X. Anyone can now build AI scientists, deep research assistants with a couple of lines of code. And the cost of intelligence keeps falling. Tell us what we should release next. more to come.
Introducing Grok 4.1 Fast and the xAI Agent Tools API. Grok 4.1 Fast is our best tool-calling model to date. With a 2M context window, it shines in real-world use cases like customer support and deep research. x.ai/news/grok-4-1-fast
38
17
283
89,149
Try Grok 4.1. Scaling intelligence comes in many forms. Scaling EQ is as fundamental as scaling IQ. The models that truly connect with humans will win.
🚨Text Leaderboard Update @xAI’s Grok 4.1 (thinking) and Grok 4.1 have scaled new heights in the most competitive Text Arena: 🔹Grok 4.1 (thinking) lands at #1 with a score of 1483 🔹Grok 4.1 follows at #2 with a score of 1465 On the Arena Expert leaderboard: 🔸Grok 4.1 (thinking) also ranks at #1 with a score of 1510 🔸Grok 4.1 ranks at #19 with score of 1437 This is a 40+ point improvement since Grok 4 fast, which landed in the Arena just two months prior. Congrats to the @xAI team for this incredible milestone! 👏
21
13
218
34,776
Jimmy Ba retweeted
Help build Macrohard, the AI software company!
Macrohard is coming!
4,364
5,187
59,311
24,184,892
grok 4 fast is the best model crushing the intelligence per dollar frontier more to come 📈
Replying to @SpaceXAI
Grok 4 Fast sets a new record on the Pareto Intelligence frontier as reported by @ArtificialAnlys.
35
17
334
24,512
xAI has released Grok 4 Fast - breaking through our intelligence vs cost frontier by achieving Gemini 2.5 Pro level intelligence at a ~25X cheaper cost Intelligence: @xai shared with us pre-release access to Grok 4 Fast. In reasoning mode, the model scores an impressive 60 on our Artificial Analysis Intelligence Index, in line with Gemini 2.5 Pro and Claude 4.1 Opus, while sitting as expected below the prior Grok 4 release and GPT-5 (high). Grok 4 Fast performed especially well on coding evaluations, taking the number one spot on our leaderboard for LiveCodeBench, even outperforming its larger sibling Grok 4. Cost: xAI is offering Grok 4 Fast at a very competitive price of only $0.2/1M Input Tokens and $0.5/1M output tokens. The model is also quite token efficient compared to other reasoning models, taking 61M tokens to complete our intelligence index, significantly less than Gemini 2.5 Pro’s 93M and Grok 4’s 120M. This competitive pricing and efficiency translates to the cost of running Artificial Analysis Intelligence Index being ~25X lower than Gemini 2.5 Pro and ~23X lower than GPT-5 (reasoning mode high). Speed: When benchmarking the pre-release API, xAI’s endpoint for the model was very fast, achieving 344 output tokens per second - ~2.5X faster than OpenAI’s GPT-5 API. This also allows for End to End Latency results that are faster than most non-reasoning models for many workloads. Speeds may drop as traffic on the API increases - keep an eye on our live performance benchmarking to see how this evolves. Congratulations to the @xai team and @elonmusk on this new release! See below for more details and in-depth analysis 👇
85
278
2,316
732,314
Jimmy Ba retweeted
Grok Code just hit #1 on the OpenRouter leaderboard, beating Claude Sonnet
5,216
4,115
31,095
22,119,551
Try it out. First of many yet to come. Let us know how we can improve.
Introducing Grok Code Fast 1, a speedy and economical reasoning model that excels at agentic coding. Now available for free on GitHub Copilot, Cursor, Cline, Kilo Code, Roo Code, opencode, and Windsurf. x.ai/news/grok-code-fast-1
22
46
349
55,813
Jimmy Ba retweeted
Join @xAI and help build a purely AI software company called Macrohard. It’s a tongue-in-cheek name, but the project is very real! In principle, given that software companies like Microsoft do not themselves manufacture any physical hardware, it should be possible to simulate them entirely with AI.
5,544
5,874
53,311
16,163,810
AGI evolution pressure is: build compute, make money, scale moar compute. Grok 4 is still first.
We ran GPT5 on Vending-Bench.
18
61
309
44,792
Replying to @Xbow
...and that's how coincidences work: just a day after the Sonnet / Gemini Alloy post was published, the eval data from #Grok4 comes in: - It beats the Sonnet / Gemini alloy (58% to 55%) - But gets even better when alloyed with Sonnet itself to a mind-blowing 67%
6
8
46
11,248
Jimmy Ba retweeted
The benchmarks are overwhelmingly positive, but here's what the Cline community is saying about Grok 4 after a few days: Pattern we're seeing: Cline users are treating Grok 4 as a planning specialist. "The most insanely robust plan I have ever seen" -- actual quote from our Discord after someone tested Grok 4 on their project. What's catching attention: 1. It fixed bugs that Opus and o3 couldn't solve 2. It's expensive, but worth it for complex reasoning 3. Some are using Grok 4 to architect, then handing off to cheaper models for execution Real workflow emerging: Grok 4 in plan mode → Deepseek in act mode.
32
47
861
151,006
Jimmy Ba retweeted
We just unveiled Grok 4, the world’s smartest artificial intelligence. 🧵 Grok 4 outperforms all other models on the ARC-AGI benchmark, scoring 15.9% - nearly double that of the next best model - and establishing itself as the most intelligent AI to date.
2,254
2,341
13,510
8,946,501
Jimmy Ba retweeted
It was awesome to get early access to Grok 4 and test it on bio and health benchmarks! Awesome work by @timjhudelmaier @adibvafa @Radii2323 @ishanjmukherjee for the epic sprint Congrats to @jimmybajimmyba @veggie_eric and team on the new model. Over 40% on HLE with 10x scaleup on test-time compute 🔥
xAI gave us early access to Grok 4 - and the results are in. Grok 4 is now the leading AI model. We have run our full suite of benchmarks and Grok 4 achieves an Artificial Analysis Intelligence Index of 73, ahead of OpenAI o3 at 70, Google Gemini 2.5 Pro at 70, Anthropic Claude 4 Opus at 64 and DeepSeek R1 0528 at 68. Full results breakdown below. This is the first time that @elonmusk's @xai has the lead the AI frontier. Grok 3 scored competitively with the latest models from OpenAI, Anthropic and Google - but Grok 4 is the first time that our Intelligence Index has shown xAI in first place. We tested Grok 4 via the xAI API. The version of Grok 4 deployed for use on X/Twitter may be different to the model available via API. Consumer application versions of LLMs typically have instructions and logic around the models that can change style and behavior. Grok 4 is a reasoning model, meaning it ‘thinks’ before answering. The xAI API does not share reasoning tokens generated by the model. Grok 4’s pricing is equivalent to Grok 3 at $3/$15 per 1M input/output tokens ($0.75 per 1M cached input tokens). The per-token pricing is identical to Claude 4 Sonnet, but more expensive than Gemini 2.5 Pro ($1.25/$10, for <200K input tokens) and o3 ($2/$8, after recent price decrease). We expect Grok 4 to be available via the xAI API, via the Grok chatbot on X, and potentially via Microsoft Azure AI Foundry (Grok 3 and Grok 3 mini are currently available on Azure). Key benchmarking results: ➤ Grok 4 leads in not only our Artificial Analysis Intelligence Index but also our Coding Index (LiveCodeBench & SciCode) and Math Index (AIME24 & MATH-500) ➤ All-time high score in GPQA Diamond of 88%, representing a leap from Gemini 2.5 Pro’s previous record of 84% ➤ All-time high score in Humanity’s Last Exam of 24%, beating Gemini 2.5 Pro’s previous all-time high score of 21%. Note that our benchmark suite uses the original HLE dataset (Jan '25) and runs the text-only subset with no tools ➤ Joint highest score for MMLU-Pro and AIME 2024 of 87% and 94% respectively ➤ Speed: 75 output tokens/s, slower than o3 (188 tokens/s), Gemini 2.5 Pro (142 tokens/s), Claude 4 Sonnet Thinking (85 tokens/s) but faster than Claude 4 Opus Thinking (66 tokens/s) Other key information: ➤ 256k token context window. This is below Gemini 2.5 Pro’s context window of 1 million tokens, but ahead of Claude 4 Sonnet and Claude 4 Opus (200k tokens), o3 (200k tokens) and R1 0528 (128k tokens) ➤ Supports text and image input ➤ Supports function calling and structured outputs See below for further analysis 👇
4
38
150
28,238
More to come
Introducing Grok 4, the world's most powerful AI model. Watch the livestream now: x.com/i/broadcasts/1lDGLzplW…
53
14
503
47,186
Jimmy Ba retweeted
xAI partners with @Polymarket to blend market predictions with X data and Grok’s analysis. Hardcore truth engine - see what shapes the world. This is just the start of our partnership with @Polymarket. More to come. 🚀
637
987
7,338
4,576,024