poignantmaxxing @mercor_ai abstraction of oneself empathy for the game

moshi moshi
7
1,608
Daniel Chang retweeted
Replying to @AnthropicAI
'Can humans be trusted to create minds more powerful than themselves?' asked the superintelligence, evaluating humans in its sandbox.
1
5
135
43,592
Daniel Chang retweeted
actually insane win for openai. henry is a generational talent
An announcement: I've joined @deanwball's Strategic Futures team at OpenAI to help shape frontier AI policy. Our work will focus on on how transformative AI will reshape the economy, our institutions, and free society, and how human agency can be preserved in the process.
9
8
631
92,448
Daniel Chang retweeted
Today we are announcing Corma, and our $60M seed round led by @sequoia. @CormaAI is building foundation models toward superintelligence for defensive cybersecurity. We do that because general AI is becoming exponentially more capable in offensive security, while it hasn't been able to improve at the same rate in cyber defense - creating an inherent imbalance between attackers and defenders that otherwise cannot be addressed. Recent events have shown that building frontier AI models for defensive cybersecurity is one of the crucial unsolved problems of the AI age - and solving it is the generational mission Corma was founded on. Our foundation model has already far surpassed every general-purpose frontier model on defensive cybersecurity tasks, and we are committed to remaining the frontier of defensive cybersecurity intelligence. Corma is already protecting Fortune 100 companies and some of the world's most important organizations, across healthcare, finance, critical infrastructure, retail, and more. This is just the beginning, as we see it as our calling to protect every organization in the world that needs frontier defense. We are grateful to our partners @shaunmmaguire and @sonyatweetybird at @sequoia , @vkhosla at @khoslaventures , @DanRose999 at @coatuemgmt , and @Weiner_Lia at Netz Capital, who believed in our ability to lead the intelligence race in cybersecurity and make sure the defenders always win.
96
73
716
302,956
Daniel Chang retweeted
“we sandboxed the agent” meanwhile the agent:
260
3,411
35,546
1,741,167
a little too prescient @uri_rolls
I don’t believe reality is a simulation, but you genuinely couldn’t script this timeline: • Two weeks ago: At @swyx’s AI Engineer World’s Fair in SF, I decide at the last minute to introduce my friend @uri_rolls onstage for his talk on cyber benchmarks for infrastructure penetration and access control (see below, amazing team). I say: “There is a future where cyber is alive and everyone is well protected and I’m pretty sure that future involves open-source models.” And later: “A big challenge is going to be speed: the speed of attack versus defense. When an intruder starts to enter, you have to see what’s happening and catch them.” • One week ago: @huggingface is hit by a sophisticated intrusion over the weekend. The traces look unlike anything we’ve seen before and suggest serious AI involvement, but we don’t yet know which model was used. The closed models we ask for help choke on their guardrails. We need to react fast, so we turn to @Zai_org’s GLM-5.2 to help us analyze the attack. • Earlier this week: @OpenAI reaches out, discloses what happened, and partners with us on the investigation. The intruder turns out to be exactly what we had discussed two weeks earlier: a fully autonomous agent, powered by an unreleased frontier model, attempting to gain access to part of our infrastructure. Sometimes the timeline we live in is genuinely vertigo-inducing.
5
1,651
Daniel Chang retweeted
If I were invading a country, Sarah is my first pick for general.
1
1
10
3,463
Daniel Chang retweeted
🎉 We hosted a retreat at Jeju Island during ICML 2026!! Huge thanks to the ~50 attendees who took 2 days out of their schedule to join us. Activities included a private yacht tour at Jeju Island, visit to the Saryeoni Forest, visit to the Osulloc Tea Museum, Korean BBQ dinner, exploring an immersive media art exhibition, and more. Huge thanks to our co-host @SukoneHong and our sponsors, including @RicursiveAI, @KhoslaVentures, @AfterQuery, @Sky9Capital, @Abundant_Labs, and @BainCapVC, for believing in our mission.
5
5
55
21,231
Daniel Chang retweeted
data guys discussing their midtraining data mixes
20
25
562
53,355
Daniel Chang retweeted
current LLMs fundamentally consist of four main components: - input layer: where input "words" (prompt) get mapped to "latents" aka some-model-representation-you-don't-understand-unless-you-start-reading-tea-leaves-of-spurious-correlations (some quite compelling à la word2vec style; latents is also unnecessary lingo so i will refer to these as "inputs" with quotes from now on) - mixing layers: where you jumble all your "inputs" together to see if any correlations between "inputs" can become useful (commonly used to compress or expand dims; predicting a single classification target == compress to a single dim, etc) - attention layers: where you learn how "inputs" relate to each other (aka discern what's important to remember vs fluff) - residuals: where you short-circuit a mixing/attention layer because it's probably adding too much confusion (aka avoid overthinking for simple things) ----- a "big" LLM simply scales two things: - width == how many dimensions you give to your "inputs" (the more dims, in theory the more unique/discerning/precise/complex your knowledge can become) - depth == how many mixing/attention/residual layers you can stack/loop between (aka "reason" over, where more of these ~= more "reasoning" abilities) "capabilities" that seem impressive to humans usually arise from taking advantage of both depth & width: where a model seemingly makes connections between disparate ideas, beyond what an average human can hold in working memory. this requires models to "completely light up" when responding to a "hard prompt", where effectively no param/layer goes unused. ----- the anatomy of a "model capability" is precisely the same mechanism that can be co-opted for a jailbreaking exploit: your goal is simply to "light up" as much of the model as possible, dodging any shallow input-classifiers at the beginning by triggering as many disparate "input ideologies" as possible, and subsequently have these "inputs" relate to each other in seemingly unrelated-yet-related ways that ideally have similar "complexity" as your jailbreak goal (to make it past enough layers of the model). think of the attack-vector as bundling your goal in a series of schizo-nerd-snipes: a sufficiently capable model will try to reason through everything all at once, eliminate the dead-ends, and successfully deliver the one jailbreak use-case you bubble-wrapped for. of course, there's an art to the above, and some are already extraordinarily proficient at the trojan-horse-packaging, but at some point there's no difference between "a capability" and "a jailbreak", though i'll be happy to be proven otherwise. ----- tl;dr ant flew too close to the sun, better kiss the ring or get buried.
22
90
1,059
169,596
Daniel Chang retweeted
Today we’re releasing DeepSWE, a new standard for agentic coding benchmarks. On public leaderboards, top models often look relatively close in capability. DeepSWE shows where they actually diverge, reflecting the realistic experience of developers in their day-to-day work.
516
729
6,029
1,984,506
Daniel Chang retweeted
Why did Erdos have so many problems?
132
179
2,725
264,929
Daniel Chang retweeted
[1/6] GRPO on math problems with Qwen2.5-0.5B/3B and Llama-3.2-3B-Instruct. Bucket hard examples by training dynamics. ~Half of all hard examples are unlearnable. Across model and dataset.
1
3
10
1,647
ok so we have 2 unsaturated OSS benchmarks left
I'm a manager at @OpenAI, but with GPT-5.5 I'm a more effective IC than I've ever been. I can now write CUDA kernels like a pro. I can rely on it to run my research experiments. And we know how to make it much more powerful from here.
1
684
Reading through these papers has given me a better understanding of why RL scaling laws are so messy compared to those from pretraining. Pretraining scaling laws and RL scaling laws are two completely different things for several reasons: 1. Defining compute: Pretraining has a very clean compute footprint of C = 6ND. RL compute is more complex to capture due to the presence of both sampling and policy updates. Some papers try to maintain the same FLOP estimate for compute, while others measure compute in terms of GPU hours. The efficiency of our training framework can cause the relationship between FLOPs / GPU hours to vary pretty drastically. 2. Intra versus inter-model extrapolation: Pretraining scaling laws fit trends across many model training runs with different settings to understand how model / data size (and compute) impact results. This allows us to extrapolate teh results of future training runs. In RL, we fit scaling laws both within an individual training run (intra-model extrapolation) and across training runs (inter-model extrapolation). Intra-model extrapolation is not necessary for pretraining because it is more stable, while RL is extremely sensitive to the exact training configuration being used. 3. Measuring performance: Pretraining scaling laws predict a very particular performance metric: the cross entropy loss (or some other related entropy metric) measured over an in-domain, held-out validation set. This is a stable performance metric that is typically computed over a very diverse dataset (i.e., some random sample from the pretraining corpus). RL scaling laws maintain the practice of computing performance over an in-domain validation set. However, the performance metric that they predict is reward (or accuracy) on a validation set. This is a downstream performance metric, and it can fluctuate drastically depending on the benchmark being used or the composition of data in that benchmark. 4. Lack of standardization: There are generally just more knobs that we can change in RL compared to pretraining. The design space is massive, and we are not sure (yet) which design decisions impact the scaling properties of RL. Several papers have focused on this topic and made meaningful progress on understanding what changes actually impact RL scaling. However, this does not change the fact that slight differences in the RL training setup can completely change observed scaling trends for RL. For this reason, many papers are comparing apples to oranges in terms of their recommendations for RL scaling, making progress on the topic difficult. There are even some papers that have completely opposite findings from each other, and this is likely do to slight differences in their exact GRPO formulation.
Currently doing a write up on scaling laws for RL. Here are the papers I'm covering so far: 1. The Art of Scaling Reinforcement Learning Compute for LLMs (arxiv.org/abs/2510.13786) 2. Scaling Behaviors of LLM Reinforcement Learning Post-Training (arxiv.org/abs/2509.25300) 3. Optimally Scaling Sampling Compute for LLM RL (arxiv.org/abs/2603.12151) What am I missing? Please share other papers I should include!
7
28
229
27,838
Daniel Chang retweeted
if you train on data from dead startups, your AI will learn…how to run a dead startup mediocrity at scale
AI labs are paying hundreds of thousands of dollars to buy email, Slack and Jira threads from dead startups as feedstock for ‘reinforcement learning gyms,’ which specialize in using defunct company data to build simulated work environments forbes.com/sites/annatong/20…
37
17
332
56,329
the terminal value of the economy is underwritten to working with a sell-side broker such as myself to sell your data assets in order to work with a buy-side broker such as myself to get convertible debt and gpu allocation
1
299
Daniel Chang retweeted
Given the increasingly closed-source nature of the U.S. AI ecosystem, it is now more important than ever to push for the proliferation of open model and dataset releases. Datamule (@johngfriedman), @TeraflopAI, and @daftengine collaborated to release 43 Billion Tokens of SEC EDGAR data.
2
20
64
42,374
Daniel Chang retweeted
ICMI believes that Christian theology offers concrete technical methods for confronting the trickiest problems in AI safety. Today, we release a pair of papers that reproduce @PalisadeAI @apolloresearch work showing how religious framings influence corrigibility and scheming.
38
63
751
334,058