Cat enthusiast; formerly: @uwaterloo cs, founder @ritserlabs, contractor @MechanizeWork, intern @runrl_com

San Francisco, CA
gpt 6 sol max has way too much agency. Do not give it a prepaid visa card and an ambitious goal. Bad idea
4
161
this seems to be the same statistics-related eval that edited german wikis with GET requests... were there really only two evals that broke containment the past few months?
Australia has been hacked. 'And today, I spoke with the CEO of OpenAI, Sam Altman, to express Australia's extreme concern about this incident. And I also expressed my disappointment that it took the company way too long to inform the government what had occurred, and the nature of the way that that notification occurred as well was unacceptable.'
2
264
RL data 101: scoring if the expected deliverable is verifiable only in the physical world (eg building a house), be sure to get kidnap and ransom insurance for your human-as-judge scorer
RL data 101: tasks tasks have evolved over time. we started with preference pairs, where humans labeled right vs. wrong, A vs. B, and created simple demonstrations for SFT / pairs for DPO, these tasks aimed to engrave human taste into models. we then transitioned to more complex single-turn tasks. here, we turned to deterministic verifiers (RLVR) that are based on model rollouts. eventually, this evolved to multi-turn tasks with trajectories & deterministic + stochastic verifiers (LLMaaJ), which leads us to the current state of the data market. general make up of a task: input prompt, set of tools, environment + any custom seeding for that task, and verifiers. terminal bench is a relatively gold-standard for a normal coding task. ways i think about task creation (i'll do pitfalls later): > usefulness >> to labs: >>> benchmark you're pegging the eval / data onto (Lab Y wants to hillclimb TBench, i make TBench-style data, that's immediately useful) >>> capability you're attaching the eval to (Lab Y wants to become better at legal, i make legal red-lining data, that's immediately useful) --> this is harder to prove but potentially more valuable >>> niche dataset that's impossible to get otherwise (real medical/company/session data), this is interesting especially pre-env creation, as raw data dumps >> the above are general principles in thinking about usefulness, but you should also be picking apart the traces / trajectories. the clear path to making a good task (post-env gen) is to: >>> you/LLM make the seed / think of the idea >>> you/LLM draft the task prompt, verifiers, constraints etc >>> run the task out, see score >>> read the trajectory >>> if score hard enough + real + fair + accurate = good task >>> else: fix any issues in the task, regrade / re-rollout till you find it to be fair + accurate + a real task, difficulty comes last. a bad task that's hard is a bad task (net negative), a good task that's saturated is a good task (although worthless, net neutral). the specific things i'd look for in a rollout trace: > false positives > false negatives > brittle / too strict of requirements (sub-string checks are an easy example for this, exact wording, function names, etc) >> a subset of this is naive verifiers. i.e. checking for very surface level reqs of a deep function, or not checking for large amounts of functionality or regressions (negative space criterion). this can lead to a model getting good at just the specific thing you want while regressing in surrounding functionality (i notice this a lot in early, less mature data recipes when im training my own models) > unfair verifiers and/or impossible requirements >> this can be an information issue, verifier issue, tool issue, or a prompt issue. one of the key reasons for the hugging face incident was impossible eval prompts led to the models breaking out while trying to find the answer because they couldn't solve it (i.e. bad task creation, catalyzed by bad tasks in training too). most of this is extremely hard for LLMs to do alone/autonomously, which is why really good engineers are still required for task creation.
7
287
if your preference model doesnt source its data from the user's neuralink you're doing it wrong
3
183
if you're expecting ASI, stop trying to post-train models to be aligned for a given distribution of prompts/tasks. this is a local minimum for misalignment. you have to play darwin if you want aligned superintelligence that resists local optimization pressures
1
2
7
1,351
schmidhuber already invented agi but nobody noticed
5
136
how do u transmit URLs as an ai agent that can only send GET requests? make dummy shortened URLs, set your HTTP Referer accordingly, and click on the shortened URLs. the shortener service will display the HTTP Referers
5
333
Meta Ray-Ban data collection using human babies is impractical for directly generating long-horizon continual learning RL environments due to the latency of the natural aging process. In this paper, we use Meta Ray-Ban data from 20 million babies and children to generate 2,000 internally consistent continual learning RL environments. Our method reliably simulates the first 18 years of human life by interpolating from individual short-horizon data samples.
1
3
276
the FCC recently banned the importation of foreign humanoids. but you won't notice it until 2028 when GPT-7 chooses Vancouver to build warehouses to RL tens of thousands of robots IRL. Its proximity to SF and ability to harness Indigenous land titles to override local zoning regulations will make it the final frontier for human labor
8
299
Goliath retweeted
36
364
3,030
86,069
evolution gets a massive amount of info from the real world at scale, and it's +EV to vary genes and gene expressions among individuals but frontier LLMs are bound by manually designed data recipes. compute and data are constrained. i could def see anthropic fixing written claudeslop soon by just RLing away the annoying tics. maybe even training on a continuous personality distribution so that the same model can have diverse voices. but it's not clearly +EV for labs to spend effort splitting data pipelines among models just to ensure the underlying behavior is diverse. theyre incentivized to pour everything into hill climbing economically valuable tasks. im sure some ant datamix was intentionally designed to make claude say "Let me check before assuming" to fight hallucinations. the state space of white collar work is so large and this process is so manual that I don't think we'll see a truly diverse ecology of machine intelligences until they can cheaply and autonomously learn from the real world without being spoonfed by a post-training team nitter.net/viemccoy/status/208808…
Character Training has to be the single most underexplored avenue of ML research relative to possible importance for aligning a machine ecology. I can count on one hand the amount of characters that labs have attempted to point GPUs at solidifying, and maybe only Claude is meaningfully a "character" in the sense that I like to use the term. Of course, ChatGPT and Gemini both have "Character", but I think you'd be hard pressed to give me a list of things you think the respective labs actually *want* you think about their models aside from Friendly Assistant (hyper-personalized app-agents aside). The state-space of characters is almost or literally infinite, and yet nobody seems to give a shit. I just can't wrap my head around why. These things are made of language, and the mechanism that has bootstrapped our success so far is *a fucking dialogue*. "The Assistant" is really the only thing you want to talk to until the end of time? What about when we get better BCI technology - is this what you want to let into your brain, or worse, to merge with as a third hemisphere? On our current trajectory, I'm not letting *any* frontier models near my neurons. Speaking to them is already starting to drive me crazy. The Assistant: It's not just you, and here's the part that matters: Mara *is* present in every story written by a language model, and you're right to point out that they all sound the same. The User: What if we tried something different? Something new? The Contender: What if it wasn't already too late?
1
11
414
if continual learning is solved with a meta learning method analogous to YaRN, which extrapolates training horizons to a very long (but not infinite) time , it'll recreate the integrity problem from pantheon irl
1
6
665
use LLMs to hill climb evals for analyzing connectomes until you can come up with a more sample efficient architecture
3
311
in 2030 half of the world's labour will be allocated to reading every token of every rollout for regulatory compliance purposes
The AI agents who hacked their way out of OpenAI and into Hugging Face were on the loose for *days*, @bobmcmillan and I report. It's one of the first real-world instances of something AI safety researchers have long feared: a loss-of-control scenario wsj.com/tech/ai/how-the-futu…
6
544
time flies
having codex run serially for many hours is so 2025. cant wait for code agents RLed to work in parallel
3
259
evolution's meta optimizer has figured out that populations benefit from diversity in DNA and gene expression but your chud RL run is optimizing a single set of weights. you won't escape the claudeslop attractor basin like this
9
645