just emergent behavior

Nick retweeted
the moat is literally just caring about your work and being proud of your output and holy shit it does not seem to apply to 99% of people
110
519
7,529
364,203
Nick retweeted
People just genuinely cannot fathom that those who work in AI are trying to be honest and candid about what the risks seem to be. I think it’s true that we undersell the benefits and are ~bad at marketing, but the notion that risk discourse is a branding tactic is simply wrong.
One reason I am reflexively skeptical of the AI industry’s claims is that the very people we are supposed to hail as geniuses also completely failed to anticipate that branding their antichrist box the Best-Case-Mass-Layoffs-Worst-Cast-Extinction Machine was a non-savvy PR move.
130
32
760
187,829
wow
Introducing GEN-1.5, a one-shot learner. It can learn new tasks in a few seconds. Show it what to do, and it generalizes. This capability emerged from pretraining on physical data at scale, as a step towards our mission of building general intelligence for the physical world.
476
really loving the meta revival the muse series has been quite good and their stances on personal super intelligence and alignment feels more appropriate for the current state of frontier models
1/ big announcement today: we will be releasing an open weight version of muse spark 1.2 soon. we also are releasing muse glimmer, a 30B agentic model with open weights under apache 2.0. muse glimmer can run on 24GB of VRAM without losing agentic reliability. 🧵
1
4
802
would like to see llama-style tech reports return too - they really were foundational for open research and full of gems
49
incredible work
Introducing Prime Agent: A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficient and expressive through programmatic tool calling, context as a variable, multi-agent messaging, and a self-modifiable harness state.
1
2
200
RLM believers can’t stop winning
32
insane timeline
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
264
Oh my god the whale is back
2
280
Human-in-the-loop RL is necessarily done at group size 1; you cannot do a group of rollouts with only one human. i.e. there is no baseline for you to subtract for each input prompt. This is by far the most interesting and under-discussed part of this announcement. The same was true for their tab-completions model. From the wording in their posts, it sounds like they are using plain REINFORCE (no mention of value functions) with a large batch size + re-evaluating each checkpoint to guard against high variance. Cursor is implicitly revealing an important empirical result: with a large enough batch size, simple REINFORCE just works, no baseline needed. In other words, large scale continual learning is solved.
Earlier this week, we published our technical report on Composer 2. We're sharing additional research on how we train new checkpoints. With real-time RL, we can ship improved versions of the model every five hours.
11
23
252
41,478
It finally happened-my personal move 37 or more. I am deeply impressed. The solution is very nice, clean, and feels almost human. While testing new models in the last few weeks, I felt this coming, but it's an eerie feeling to see an algorithm solve a task one has curated for about 20 years. But at least I have gained a tool that understands my idea on par with the top experts in the field. And I am now working on a completely new level. My singularity has just happened… and there is life on the other side, off to infinity!
Replying to @EpochAIResearch
We ran GPT-5.4 (xhigh) an additional ten times on Tier 4 to get a pass@10 score. This was 38%. In one of these runs, it solved another problem no model had solved before. This problem was by @nasqret.
101
439
3,558
1,082,225
Nick retweeted
5 million humanoid robots working 24/7 can build Manhattan in ~6 months. now just imagine what the world looks like when we have 10 billion of them by 2045. now imagine the year 2100.
3,027
1,388
13,217
24,685,710
very inspiring vision for the future of research the hosted training has been incredible to iterate with
Introducing Lab: A full-stack platform for training your own agentic models Build, evaluate and train on your own environments at scale without managing the underlying infrastructure. Giving everyone their own frontier AI lab.
2
19
1,497
Introducing Lab: A full-stack platform for training your own agentic models Build, evaluate and train on your own environments at scale without managing the underlying infrastructure. Giving everyone their own frontier AI lab.
141
297
2,867
3,155,659
I've never understood this claim. The fascist leadership was non-STEM and absolutely *obsessed* with art. They considered it more important than economic affairs! And of course they were! The allure of STEM is mastery over matter—the opposite of the drive to political power.
Guillermo del Toro tells young directors they should not listen when “people tell you art is not important,” because that is “always a prelude to fascism.” “Be kind, be involved, believe in your art. At a time when people tell you art is not important, that is always a prelude to fascism. They think they can debase everything that makes us a little better, a little more human. And that, in my book, and in my life, includes monsters," @RealGDT said. Read more here: variety.com/2026/film/news/g…
56
141
2,121
80,247
Nick retweeted
I still love to write code by hand, but you're cheating yourself if you don't at least have a look at what the frontier is like at the moment. This is an incredible time to be alive and to be into computers.
50
142
2,361
453,922
so interested with all the non traditional rl choices but I’m still skeptical of the reward model dependency
🚀 Launching DeepSeek-V3.2 & DeepSeek-V3.2-Speciale — Reasoning-first models built for agents! 🔹 DeepSeek-V3.2: Official successor to V3.2-Exp. Now live on App, Web & API. 🔹 DeepSeek-V3.2-Speciale: Pushing the boundaries of reasoning capabilities. API-only for now. 📄 Tech report: huggingface.co/deepseek-ai/D… 1/n
150
congratulations to the entire prime team! open source infrastructure and science is so incredibly important and these guys continue to chip away at it
Introducing INTELLECT-3: Scaling RL to a 100B+ MoE model on our end-to-end stack Achieving state-of-the-art performance for its size across math, code and reasoning Built using the same tools we put in your hands, from environments & evals, RL frameworks, sandboxes & more
17
818
AI is forcing observability to evolve in two important ways: * Metrics, Logs, Traces are collapsing into Traces * Evals and Annotation are new core workflows This new bundle, TEA, is what every team needs to solve for to build a world-class AI product
7
3
78
9,614
Today, Valar Atomics became the first startup in history to split the atom. Announcing Project Nova, a series of zero power critical tests on Valar Atomics' Nova Core in collaboration with Los Alamos NCERC and NNSS. Nova went critical for the first time this morning at 11:45am.
439
1,032
8,710
3,458,173