Computer scientist obsessed with markets, Founder builder @the_nof1

Manhattan, NY
1/6 🧵 How do you train a language model to become a self-adaptive, long-horizon optimization engine for executable trading policy search? In our new paper, "The Time Value of Evolution," we introduce Lineage-Value Policy Gradients, or LVPG, a new long-horizon RL technique for adaptive evolutionary search.
3
4
30
12,272
New work on programmable cellular automata by @Amidos2006 @utheprodigyn @togelius and myself!
Cellular automata are powerful but painful to author. Rules are a lookup table; neural CAs are more expressive but a black box. What if each CA rule was just code that you can read? Introducing Programmable Cellular Automata (PCA)
1
2
5
929
Matthew Siper retweeted
.@Amidos2006 presenting our work (led by @MatthewSiper ) on Procedural Content Metageneration via Continual Abstraction Discovery here at @ieee_cog
1
2
6
845
Matthew Siper retweeted
This week on The Information Bottleneck we're talking with Julian Togelius (@togelius ) 🥳🥳🥳 Julian directs the Game Innovation Lab at NYU and co-founded modl.ai. For years he's worked on AI and games from both directions: using AI to generate game content and test games, and using games to actually measure what AI can do. Lately he's been digging into why LLMs that write code so well still fail badly at playing games, and what that gap tells us about spatial reasoning and learning in current models. We'll get into game playing as a benchmark, procedural content generation, open-endedness, and where he thinks the AGI conversation goes wrong. What should we ask him? Drop your questions in the comments
3
13
2,409
Wild
trained on prime intellect lab :) this is nuts my mind is kinda blown
1
2
311
🧵 1/5 What happens when a language model doesn't just write code, but dynamically creates its own expert API mid-run? Instead of evolving individual game levels, our new paper evolves entire generator programs using an LLM as the mutation and crossover operator.
1
3
23
6,233
4/5 CAD raised mean final-best in all comparisons across Sokoban, Zelda, Lode Runner, and Dangerous Dave. The endpoint favored CAD whether search began with no helper library or a fixed hand-written expert API.
1
2
2
501
5/5 CAD turns program search into a process whose vocabulary changes during the run. Later programs inherit callable abstractions discovered from earlier high-fitness programs. Accepted at IEEE CoG 2026. With @Amidos2006 @togelius. Website: matt-quant-heads-io.github.i… Paper: arxiv.org/pdf/2608.17947
1
8
587
Matthew Siper retweeted
🚨 new preprint: Information Abundance Paradox TL;DR: Longer context can weaken parametric learning in LLMs. We propose the Information Abundance Paradox: when relevant information is abundant in the training context, the model has less incentive to internalize that information in its parameters.
5
30
97
19,583
Matthew Siper retweeted
At Nof1 we believe the next breakthrough in AI after reasoning is adaptation If reasoning is the ability to detect patterns and leverage them, adaptation is anticipating how those patterns will change This is a crucial capability for real world AI, and current LLMs struggle with it. We've run thousands of live market experiments and no amount of context, prompting, or harness engineering has worked. So, we've been pushing our research deeper down the stack In this paper, we trained models that can better adapt in dynamic environments like markets. They learn to value an action by the futures it opens up, not just its immediate payoff It's an early step, but we're building towards forever learners that are default forward-looking, rather than static and backward looking
1/6 🧵 How do you train a language model to become a self-adaptive, long-horizon optimization engine for executable trading policy search? In our new paper, "The Time Value of Evolution," we introduce Lineage-Value Policy Gradients, or LVPG, a new long-horizon RL technique for adaptive evolutionary search.
5
7
66
12,633
1/6 🧵 How do you train a language model to become a self-adaptive, long-horizon optimization engine for executable trading policy search? In our new paper, "The Time Value of Evolution," we introduce Lineage-Value Policy Gradients, or LVPG, a new long-horizon RL technique for adaptive evolutionary search.
3
4
30
12,272
5/6 Stronger finite-budget search: Across 90 matched paired runs evolving quantitative trading policies, LVPG increased mean sealed-test Sharpe from 0.86 to 1.32, while peak sealed-test Sharpe reached 4.13.
1
1
291
1/9 What if natural language were not just a prompt, but the genome of an evolutionary system? In “ELMER: Evolutionary Language Model that Explores and Refines,” Ahmed Khalifa, Julian Togelius, and I trained an 8B model to evolve executable trading policies in natural language. 🧵
7
4
30
6,512
8/9 Requested mutation strengths systematically alter the composition of semantic edits, while language mutations preserve more parent fitness at matched small-to-moderate behavioral displacements.
1
1
172