C and Assembly • Grok • @a_g_e_n_c //: 👾

Neo Tokyo
Pinned Tweet
I compiled Quake III's fast inverse square root on 64-bit Linux with gcc -O1 and it returned negative numbers. The long cast reads past the float into the stack, and the low bit of the return address ends up in the sign bit.
12
25
208
9,815
Per year of existence, SpaceXAI has covered far more ground than OpenAI or Anthropic. Twice the speed, and now the most compute in hand. Not planned. Installed. Colossus 1 + Colossus 2 today: 150k H100, 50k H200, 140k GB200, 440k GB300. That is 1.68M H100-equivalents. Next week: +220k GB300 → 2.23M November: +220k GB300 → 2.79M December: +220k GB300 → 3.34M Confirmed sites, same standard for everyone (Epoch AI's data-center database, Sep 25, 2026): SpaceXAI 1.7M · Anthropic 0.7M · OpenAI 0.5M That compute supports a compute-optimal model of 7.6T parameters today and 10.7T with the December tranche (120-day run, 40% MFU, Chinchilla scaling). Stated plan from the Q2 investor call: over 2 GW by end of 2026, "closer to 10 GW than 5 GW" in 2027, up to 20 GW of infrastructure in the pipeline, and an FCC filing for up to one million orbital data-center satellites. And they got here faster than anyone. Time from first scored model to first frontier model at ECI 150 or above: Grok 4.20, Feb 2026: 14 months Claude Opus 4.5, Nov 2025: 28 months GPT-5, Aug 2025: 29 months Never bet against Elon. Chip counts: Elon Musk, Sep 25, 2026. Site data: Epoch AI. Conversion: Epoch's factors, GB200/GB300 = 2.53 H100e. Capability: Epoch Capabilities Index, Sep 25, 2026.
Replying to @minchoi
Colossus 1 is 150k H100, 50k H200 and 30k GB200. Colossus 2 is 110k GB200 and 440k GB300. Another 220k GB300 will be fully operational next week and another 220k in November. If we get lucky, yet another 220k GB300 by late December.
1
14
1,477
x86 has had its own inverse square root instruction, rsqrtss, since the Pentium III in 1999. On my Threadripper, rsqrtss plus one Newton step runs as fast as Quake's function with about 8,000 times less error.
I compiled Quake III's fast inverse square root on 64-bit Linux with gcc -O1 and it returned negative numbers. The long cast reads past the float into the stack, and the low bit of the return address ends up in the sign bit.
3
1
21
2,542
gdb shows the long cast reading past the float into the stack, and the low bit of the return address landing in the sign bit.
I compiled Quake III's fast inverse square root on 64-bit Linux with gcc -O1 and it returned negative numbers. The long cast reads past the float into the stack, and the low bit of the return address ends up in the sign bit.
1
3
16
2,654
Update. It draws now. I had 128K words of ROM to work with, two 1989 EPROMs, and 48K of RAM. It took ~80 training runs to get a transformer that fits and still draws something. The weights are 4 bits each, 4 to a word. They index a learned codebook. The CPU has no SIMD, so a table lookup costs the same as unpacking a nibble anyway, and the codebook version lost 0.01 bits per pixel, where plain int4 lost 0.13. 2 and 3 bits both came out worse, even after spending the saved ROM on more layers. The cache of the last 256 pixels is int8, with 2 KV heads shared by 8 query heads. Cutting the cache 4x freed enough ROM for a third layer. I also added three builtins to the programming language for streaming dot products. About 10 cycles per weight instead of 40. A picture takes 30 seconds now instead of 3 minutes. 308,224 params, 94K ROM words, 41K RAM words, 6 billion cycles per picture. Not released yet. Putting this on an FPGA, releasing a video of it drawing off a coin cell, and open sourcing it on Hackaday.
Building a computer from scratch with Grok 4.7. Spec-sheets created with Opus 5.5. In folklore, a golem is made of clay and comes alive when you write אמת ("truth") on its forehead. The CPU is the Golem-16. One instruction format, 8 registers, and a 40-bit MAC, the same trick 80s DSP chips used. Clay is a typeless language in the spirit of B, the language that came before C. Every value is one 16-bit word. Clay compiles to Golem assembly. The SI model is Emet-107K. 2 layers, 4 attention heads. Most weights are 8-bit integers, and every rescale is a single bit-shift instruction. It fits on two 1980s-sized EPROMs. At 10 MHz it should speak about 10 characters a second. You'll be able to watch registers flip as it thinks.
7
10
67
4,584
tetsuo retweeted
Legacy media REFUSED to show you a SINGLE SECOND of President Trump's speech and historic state dinner with Xi Jinping tonight at the White House So here's the ENTIRE thing, from start to finish, with translations for Xi's speech DO NOT LET LEGACY MEDIA KEEP YOU IN THE DARK
1,522
20,910
71,238
3,888,027
tetsuo retweeted

Tesla Semi

Semi Rollout

119
579
4,235
184,205
Starlink enables access to truth, choice, and freedom of thought and speech.
🚨 ALERT: At the UN, Benjamin Netanyahu held up a Starlink device and took aim at Iran’s internet ban — then said leave it for the Iranian delegation so when they defect, they can finally tell the truth. Freedom of speech.
27
13
264
10,738
The most important equation in statistics and the foundation of linear regression. Professor Gilbert Strang, MIT.
10
33
278
12,657
These are insane. 100 grams on your face. The rest of the headset lives on a fiber tether in your pocket. Meta VR Glasses. 5K micro-OLED. 37 PPD. 120Hz. Gaze + pinch. $1,300. Spring 2027. 70° FOV.
85
43
711
94,806
Project Based Tutorials in C. Everything from computer architecture to game dev to OS internals, all through hands-on builds. Good entry point if you're learning the language and want something concrete to work toward. github .com/7etsuo/project-based-tutorials-in-c
14
115
847
24,574
Building a computer from scratch with Grok 4.7. Spec-sheets created with Opus 5.5. In folklore, a golem is made of clay and comes alive when you write אמת ("truth") on its forehead. The CPU is the Golem-16. One instruction format, 8 registers, and a 40-bit MAC, the same trick 80s DSP chips used. Clay is a typeless language in the spirit of B, the language that came before C. Every value is one 16-bit word. Clay compiles to Golem assembly. The SI model is Emet-107K. 2 layers, 4 attention heads. Most weights are 8-bit integers, and every rescale is a single bit-shift instruction. It fits on two 1980s-sized EPROMs. At 10 MHz it should speak about 10 characters a second. You'll be able to watch registers flip as it thinks.
23
34
326
13,962
Logic and microcode behind modern CPUs.
33
171
1,755
38,087
tetsuo retweeted
Interesting. Grok 4.7 is performing fairly well for a smallish model.
We ran fresh Next.js evals. The tally: ① Opus 5.5 [𝟿𝟽%] ② GPT 6 Sol [𝟿𝟽%] ③ Fable 5.1 [𝟿𝟽%] ④ Grok 4.7 [𝟿𝟺%] Notably, Grok is 2x-7x cheaper
1,822
1,443
12,710
5,705,513
SUPER INTELLIGENCE
CSPAN
118
113
809
37,560
Grok 4.7 is my favorite coding model right now. I highly recommend Grok Heavy for token maxing. Grok 4.7 built a custom game engine and playable FPS in C, in about 7 hours. It uses SDL2 and OpenGL, with a fixed 72 Hz sim seperated from rendering. The renderer uses GLSL shaders, OBJ meshes, and triplanar world texturing.
50
40
457
22,946
19
27
240
11,047
We live in a sci fi novel and this is the support ticket. “A datacenter incident is affecting all Grok models. The team is investigating.” - Grok Build
18
12
284
21,437