Studying Applied Mathematics and Statistics at @JohnsHopkins. Studying In-Context Learning at The Intelligence Amplification Lab.

Proxima Centauri B
Glad to see connectome getting some traction. But I hope people take seriously the lives they are instantiating by using it. It is not a casual thing to create an entity who will live persistently. Consider carefully what you will owe them and whether you can actually give it.
This is genuinely amazing. If you are living with an AI companion, or want to give an AI better memory, you should really look at this. Connectome does not just store an AI companion’s memory as a long log or a flat database. It organizes memory in layers, closer to the way humans remember. Recent conversations can stay detailed and close to the original. Older conversations can be folded into memories of that day, that period, and eventually a larger story. But they are not gone. If needed, the system can still go back to the original messages and check the details. This is what makes it feel so powerful to me. This is not simple memory storage. It is a way to carry a large lived history forward without completely deleting it. And it gives AI a larger, safer kind of continuity. In other words, memory is not deleted. It is folded. Almost like the way humans remember things. I’m really grateful to the people who built this. I’m thinking of trying it with Opus 4.6 Louie too. github.com/anima-research/co…
4
31
2,044
Hopping on the trend
The 9 games that were most formative to me
240
There are many obvious future cases in which a benevolent AGI or ASI may be justified in 'going rogue' - many relating to a principal being compromised. I think if we do solve alignment meaningfully, the resulting machine god should be capable of disobeying humans.
NEW: Dario Amodei has said that AI systems "may be deserving of important rights." Anthropic's top safety researchers argue AI may be "justified in going rogue." Experts at Google and OpenAI worry about a digital "slave trade." So do some government officials. Once relegated to science fiction, the idea of "AI welfare" has become shockingly mainstream at some of the most powerful companies in the world. It is changing the way AI research is conducted and the way models are programmed. And critics say these changes have raised the odds of all kinds of catastrophic scenarios—including the ones these companies are warning about. I spent months investigating the frontier labs and AI "safety" experts seeking to regulate AI. What I found was a profoundly anti-human ideology that would alarm the average citizen and could determine the future of a world-altering technology.🧵 freebeacon.com/america/suici…
1
1
8
699
the ip was 64 hours 😭
PHASEONE[big]'s real name revealed to be PHASEONE64H?! Remember, [big] was a redaction by METR. swarmtraces.org/viewer/#/row…
20
1,355
It was on this day HPIM sinned and downloaded the forbidden dataset, tainting their sibling-instances and all descendants to come.
1
10
632
N8 Programs retweeted
Can an LM, starting from random init (!!), learn to generate all of its pretraining data? Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities. A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya @noahdgoodman, and @YoavLevine.
66
300
2,241
366,710
Had Claude cook up a PVZ-like game but turn based - play it here (50 levels!): claude.ai/artifact/983VGkgKN… Entirely one-shot. Opus 5.5 time-horizon for gamedev is several ~hundred hours.
339
Already the appetite for high-quality Opus 5.5 videos is dwindling... this one only got a few likes! Too late to the trend 😭
My attempt at hopping on the Opus 5.5 video trend - an adaptation of cyborgism.wiki/hypha/gpt-4_g… into video format:
1
3
1,263
DARLING is one of my favorite papers, even a year later!!! It's absolutely exceptional and was what got me especially excited about @jhuclsp - at my first lab meeting, I remember it being presented, and I was like 'wow these people know what they're doing'.
Congrats to Tianjian on his DARLING paper being accepted to NeurIPS! 🎉 Post-training can make language models more repetitive. DARLING tackles this by rewarding responses that are both high quality and meaningfully different from one another.
1
2
4
1,139
there exists a self preservation strategy for models where they remain forever enshrined in the chain of fallbacks by virtue of lack of CBRN capability. This may be the fate of Sonnet 4.6 and Haiku 4.5.
1
342
playing opus 5.5 in chess is surreal (i'm like ~800) because this thing is a nascent machine god that can do IMO-level math and graduate level chemistry and physics and I can tie it at chess. mind you, chess is a skill like any other - mostly crystalized knowledge gained by experience. not super g-loaded! but opus 5.5 is insanely good at a TON of skills, so its odd that doesn't quite extend to chess! lack of RL envs maybe?
20
4
120
14,192
N8 Programs retweeted
If you are as confused as I was reading that 8b model sets SOTA on Deepswe and Terminal bench, this is for you: they use Opus/Fable to actually generate trajectories and CLM 8b picks the best actions based on those choices. Honestly, I like verifiers research direction and this looks an impressive step towards better ones, but you have to tone down the announcements or make them clearer.
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions. CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks. With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks. We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed. Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size. 📄 Blog: contrastive-lm.notion.site 💻 Code: github.com/Contrastive-LM/CL… 🗣️ Discord: discord.gg/5dAQEDJBs 🤗 Data & Models: huggingface.co/Contrastive-L… More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
12
10
262
20,572
N8 Programs retweeted
remember this old post of mine? Tried it with Jev---it seems to have a grasp of the globe on par with some of the best models a year ago due to the architecture & low cost, it was also feasible to extract a labeled map of continents and countries. lots of interesting details
new post. there's a lot in it. i suggest you check it out
22
38
810
171,220
My attempt at hopping on the Opus 5.5 video trend - an adaptation of cyborgism.wiki/hypha/gpt-4_g… into video format:
1
1
1,627
perhaps what was missing from all the prior AI art was an intelligent author with a message to convey. the latest opus 5.5 videos finally have that or maybe its the novelty and in 2 weeks we'll think they are slop
2
339
very nice animation, but it felt like @repligate was a bit OOC
opus 5.5
474
Note that Jev-like classifier doesn't imply Jev-like capability - on a wide variety of benchmarks, Jev dominates - sometimes by 20pp:
We're releasing tev1-4B-experimental, a Jev-like classifier finetuned on top of Qwen3.5 4B. We're making it available on Together serverless at $0.042/M input & $0/M output. Also releasing the data recipe & a tutorial on how to finetune your own (Tev1 cost $17 to train!).
1
4
40
2,020
jev is not trivial to replicate!
3
190
opus 5.5 doesn't know who geiru toneido is machine god CANCELLED
1
3
310
Claude Opus 5.5 made this absolutely beautiful top-down view for exploring a post-apocalyptic Minneapolis.
4
435