sounds to me like 1000 watts is a good step then
1
98
Agreed, It is an admirable achievement, and the article presents a measured interpretation of it. I was gently poking fun at the headline which was more in response to how Space X presented timelines for full fledged datacenters in space as part of their IPO rather than how Google has represented their work in this area.
2
82
how much compute is currently in space?
1
157
I would suspect much of the compute capacity sitting in space is on classified satellites, but outside of that StarCloud-1 put an H100 in space in 2025 which I believe at the time was the most GPU compute in space.
1
1
149
A few thoughts on why @AnthropicAI's Opus 5 sometimes felt lazy compared to @OpenAI 's models. I like to use chess as a benchmark/eval for AI models. It is a crude/toy benchmark, but it does a good job of giving a sense of when the frontier model's make leaps forward and helps temper some of the benchmaxxing. 1/
1
120
While regardless of chess performance Anthropic models have continued to get stronger in other domains, my own anecdotal experience was that Opus 4.5 and Opus 4.6 were the last Opus models to feel consistent. Ever since Opus 4.7 I feel like I am constantly having to correct and double-check the model's work. One moment the model feels smart, the next it is apologizing to me for the umpteenth time for having made a mistake. These days anytime I use an Anthropic model for any serious work I am constantly farming out adversarial reviews to codex/OpenAI models. This change in anecdotal vibes for me coincided with Anthropic's shift to "adaptive thinking". The older model's used the now deprecated "extended thinking" feature. If we add in data on the older model's we see a profile far closer to the OpenAI models as far as thinking token allocation. Adaptive thinking definitely saves money on a per API call basis, but I do wonder to what extent that comes at the cost of quality. 8/
1
33
Comparing thinking tokens across models/providers is fraught, and of limited value. Different models, different backend implementations, different tokenizers, etc. Stronger models should be able to solve more complex problems using less thinking tokens/test time compute. The older pre-adaptive thinking Opus models are weaker so even with higher thinking token allocation their performance on chess doesn't improve enough to catchup to latest OpenAI models. With that said, I do think there is value in tracking reasoning/thinking token allocation to see how often low/zero thinking token allocation corresponds to poor performance.
14
Based on the tokenizer data the new stealth model "space-bunny-alpha" appears to likely be part of MiniMax lineage.
1
1
238
Out of curiosity, I ran it through my standard chess eval gauntlet on low, medium, high, and max. It lost every game except one. Here is one of its losses, mate in 7 moves with stockfish at 1320. All match replays have been added to the library: codeandstream.com/chess/#g=s…
54
OpsConfig retweeted
Annie Dillard has died. In our Summer 1988 Issue, she wrote about the "surprising letters" and gifts that fans and readers sent her over the years—a fossil fish, a marriage proposal, a check for $2.50, and, again and again, questions about God. yalereview.org/article/annie…
1
58
257
7,402
❤️‍🩹
7
15
161
5,595
OpsConfig retweeted
Two heroes of today's Navier-Stokes story are not getting enough credit. Córdoba and Martínez-Zoroa came up with the crucial strategy. Others built on it and, with the help of AI, provided the computations that finished the problem. Thankfully @QuantaMagazine gets this right:
21
338
1,701
89,803
Two things to distinguish: Did any human or agent look at user data as part of the Navier Stokes effort? No. Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.
“we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” i mean props to them for straight coming clean. (so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan) so i’ll now give a bit on my thinking here. i actually woulda been pumped to collaborate on this, there are a lot of people at oai i like (ok, clearly some were indirectly dicks to me because of being part of the whole situation, but im a big boy, i still like them), idgaf about authorship on that step anyway, coulda been me Tristan and every fte at oai for all i care (on that Tristan would disagree:p). but on hearing the loud convo in the hallway, especially the part where a millennium prize was offered if i’d just be removed from the paper, it was kinda clear the die had been cast and things were locked. pretty wacky, unstrategic, and unnecessary, since on my side things were mostly me and claude having a good time yoloing random stuff in the corner rather than anything institutional. i also like the idea of the labs cooperating, and even better on scientific progress. it’s a shame!
556
120
2,277
668,296
Just to make sure I understand are saying you aren't sure if they had these settings turned on/off, which seems quite reasonable? Or are you saying even if these settings are turned off there is still some gray area?
25
17
582
32,347