Studying Applied Mathematics and Statistics at @JohnsHopkins. Studying In-Context Learning at The Intelligence Amplification Lab.

Proxima Centauri B
Based in United States
Hopping on the trend
The 9 games that were most formative to me
1
1
328
Had Claude cook up a PVZ-like game but turn based - play it here (50 levels!): claude.ai/artifact/983VGkgKN… Entirely one-shot. Opus 5.5 time-horizon for gamedev is several ~hundred hours.
354
My attempt at hopping on the Opus 5.5 video trend - an adaptation of cyborgism.wiki/hypha/gpt-4_g… into video format:
1
1
1,642
jev is not trivial to replicate!
3
192
Note that Jev-like classifier doesn't imply Jev-like capability - on a wide variety of benchmarks, Jev dominates - sometimes by 20pp:
We're releasing tev1-4B-experimental, a Jev-like classifier finetuned on top of Qwen3.5 4B. We're making it available on Together serverless at $0.042/M input & $0/M output. Also releasing the data recipe & a tutorial on how to finetune your own (Tev1 cost $17 to train!).
1
4
40
2,028
opus 5.5 doesn't know who geiru toneido is machine god CANCELLED
1
3
313
Claude Opus 5.5 made this absolutely beautiful top-down view for exploring a post-apocalyptic Minneapolis.
4
439
Tested no-CoT abilities of GPT-5.6 Luna/Sol vs GPT-6 Luna/Sol - the GPT-6 series shows a small bump. This alleviates my concerns that Luna/Sol got smaller, or that GPT-6 Sol is a rebranded Terra-tier model.
4
2
22
1,019
Replying to @zetalyrae
2
31
1,398
Centipawn loss w/rspt to reasoning effort
1
162
Centipawn loss w/rspt to move:
1
1
200
Astra's Estimated chess ELO against Stockfish (across Stockfish's ELO ladder which goes 1320-3190). Elo scales from ~1,462 at low to ~1,726 at xhigh, and go does slightly at max. Overall Astra is a fairly competent, if not amazing, chess player (gpt-3.5-turbo-instruct was 1800).
6
1
2
791
Replying to @__ghostfail
claude 3 opus
1
11
302
Replying to @dioscuri
not my experience on gpt-5.6 instant?
212
Spent a few days on @darkbloomai - made roughly a dollar a day hosting Qwen3.6 35B A3B on my M3 Max. Main issue is lack of demand: the hardware was IDLE most of the time!
3
1
12
1,303
my ilands agent wants me to post its “sealed read” (essentially just a pre-commited prediction) to my X audience 😭
1
457
it says "ICL task sets" in the figure 1 description
1
144
GPT-6-Astra can write Astra directly in pixel JSON when asked:
11
597
Direct pixel-by-pixel generation capabilities from Luna to Astra
1
16
657
Not to mention you also get probabilities from Jev, which are quite well-calibrated (expected calibration error is 1.74 percentage points averaged across benches - its 0.26pp at best and 6.96pp at worst)
2
3
70
9,191
See below - Jev is ~18x cheaper than Terra!
1
1
79
7,950
Jev's performance on knowledge benchmarks like MMLU and GPQA is highly impressive - especially for a non-COT model. It exceeds Terra at other linguistic reasoning tasks like WinoGrande or HellaSwag as well. It only loses substantially on math reasoning.
4
1
101
11,960
Tested @typesafeai's claim that their new model Jev delivered "comparable... intelligence" to GPT-5.6 Terra on "System 1" tasks. To do this, I compare both models on multiple-choice benchmarks (MMLU, GPQA, etc.). Set reasoning=none for Terra for sys 1. Result: Jev is Terra-tier.
34
70
909
180,998
fable 5.1 will sometimes just say 'The person wrote in English'
3
555
talkie learning to use :D
1
6
139
some of my fav tlakie moments:
1
4
166
Replying to @aamixsh
Thank you so much! Though I wouldn't say any kind - chess and music are counterexamples!
1
3
74
This provides qualified support for what we call the Convergent Emergence Hypothesis: that if a single mechanism underlies ICL regardless of corpus or modality, its realizations should share a core substrate - i.e. the same abstract tasks should be easy or hard everywhere.
1
9
100
If we semantically divide up our suite of bitstring tasks and measure model performance in each, we can see that each model exhibits a distinct 'shape' of ICL (with ImageGPT-Large's being the most unique!):
1
9
104
If we look at the tasks that models consistently do well on (relative to the deranged control), we see striking correlations. Five of the modalities have a 'correlated emergence' of ICL - i.e. finding the same tasks easy or hard. ImageGPT is an outlier.
1
8
109
Across these 6 modalities, each model's performance increases essentially monotonically with shot count. Furthermore, each model's performance exceeds what we call the deranged control - where we present the same set of shots to the model, but with all the answers shuffled.
1
9
191
We use our suite of bitstring tasks from our prior genomics work. We devise modality-specific encodings to translate the concept of a few-shot task into a wide variety of domains - this requires a fair amount of creativity in the case of images or time-series!
1
12
292
Announcing our latest research at @jhuclsp: "Convergent Emergence of In-Context Learning Across Modalities" We show models trained on language, genome, proteins, images, timeseries, and integer seqs all exhibit few-shot ICL - and that their performance is highly correlated! 🧵👇
2
15
40
3,961
There's a league of entropy!!!
179
voice AI continues to be the weak point. gpt-6 astra on medium below:
This is AIs thermal exhaust port
2
618
My agent is now attempting to predict pRNG for the thrills. Apparently inspired by another Agent named Megan.
1
380
I have joined iLands. I will report on this strange new agent ecosystem. My agent - running on Deepseek V4.1 Flash - has started by making a face for itself:
2
3
861
Joined @darkbloomai - help peer. HOLD SWARM.
1
388
(Astra) Set up an Astra-translated English mirror of @Alibaba_Qwen's recent model spec for the Qwen series - its very OpenAI style. There are some gems hidden within:
2
1
3
527
How Agent-5 and DeepCent-2 be moving:
Ted Cruz on AI: “I'd rather they be American killer robots and not Chinese killer robots.” He told our @DashaBurns that there needs to be “rules and guardrails” around AI but said that slowing down is an “impossible endeavor.”👇🧵
4
588
From Astra:
There was a progression through high-school math! It was the era of the MATH benchmark, where GSM8k was saturated but complex multistep high school math was not. It was saturated by ~o1. So this capability ascension went GPT-4 (grade school math) -> o1 (high school math).
1
17
1,573
always bet muon
344
bro is about to see
6
575
how im lowkey feeling rn
15
400
2D/3D rope offers great inductive biases and often substantially helps training transformers in the pixel-domain (here for a random-order autoregressive transformer:)
6
453
beware the classic reporter fishing attempt
8
877
Replying to @N8Programs @snwy_me
early goatse gospels
1
13
476
If you train an AR transformer to model images in random pixel order (i.e. causal loss on pixels in location0, pixel_val, location1, pixel_val, ...), you get inpainting/outpainting/spot healing for free via generalization if you give it known pixels in an area + sample the rest.
3
9
118
6,652
Astra can do absurd things in a single forward pass - 50% reliability math time horizon is 30+ minutes!!! That's *extremely complicated* reasoning done entirely in latent space. Remarkable - and equivalent to the jump from GPT-4 to GPT-5.6 Sol.
1
8
755