Leading research at @arcee_ai | Formerly Data Research Lead @DbrxMosaicAI | Visiting Researcher @meta | Ph.D | #TXSTFOOTBALL fan | linktr.ee/code_star

Redwood City, CA
I said what I said
Replying to @brianluidog
I know him personally and he is one of the sweetest, kindest human beings I have ever met. He was 9 and just trying to do his best. Also when I get drinks with him and my straw gets soggy I get to hit him with “that’s the last straw”
2
22
1,088
The people with eyes
there are NP Hard problems everywhere for those with eyes to see them
2
1
25
1,260
Replying to @zndx
3
116
This is really cool and all. I really wish it was atleast partially open to public interaction instead of locked behind a waitlist, but this FAQ on the site is kind of funny. Ok, so it's just parallel 100% perfect structured output, but not hallucination free. I honestly can't remember the last time I actually experienced agent getting structured output wrong. I know it happens, but now the agent just debugs what it did and fixes it in a harness and I don't deal with it. If the json is invalid, fails in the sandbox, and no user is around to see, does it make a sound?
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
7
2
30
3,466
Why aren't you training your models like this?
7
24
932
The answer to every challenge lies in your heart.
17
2,296
I won’t out the guilty parties, but slack has been getting out of control recently.
1
1
23
838
TXST right now
Manning is good
6
1,375
Manning is good
7
1,262
That's a lot of bookmarks
it turns out "Consider setting torch.set_float32_matmul_precision('high') for better performance." was worth paying attention to
1
1
28
2,668
Replying to @CharlesDardaman
Once all the cache gets aligned
4
102
Enough to feed the “family”
WATCH: Thieves steal 105 kg of chicken from a speeding truck in Cairo in a bizarre Fast & Furious-style heist.
3
14
1,588
What?
TIME’s new cover: Announcing the 2026 TIME100 AI, the world's most influential people in artificial intelligence time.com/collection/time100-…
11
1
48
9,110
The new @Alibaba_Qwen tech report is one of the best I have seen in a long time. You can really see the level of attention to detail to go back and challenge their own long held assumptions. the N-gram thing is super awesome too. It will be really interesting to see if this becomes standard over the next year or so.
1
5
151
10,856
The goose was not valued
Though it seems next to no one had appreciated this, US GDP simply doesn't count Nvidia's value-add! Taiwan's stats don't either (which they shouldn't), so the world's largest firm is just left out of GWP! This is largely because of a politicized decision over a decade ago. People worried that counting the sort of design work firms like Nvidia do as "manufacturing" would interfere with efforts to promote "real" US manufacturing, and it ended up just not being counted anywhere.
1
24
1,340
When I tune the dp replicate degree myself without the pretraining team's help
1
2
32
1,270
Drop the neural, it's cleaner
a neural network but it’s just the network
1
22
5,135
When you train models with big ass weights you need heavy ass riffs
2
12
598
Replying to @andersonbcdefg

ALT breaking bad oops GIF

1
34
I fixed it. Or rather, nac did github.com/arcee-ai/nac It turns out bad regex wasn't the only thing wrong in multiPL-E. Full blog coming next week, but this was too good to sit on. (1/5)
Looking at the data is so funny because in like 10 seconds you can find mistakes that people have just been ignoring for years. like this clearly bad regex replace in multiPL-e MBPP, replacing py with rs, resulting in this hilarious instruction to write a "rsthon" function
4
6
41
8,203
Most eval infra
an unspoken rule among eval creators is to come up with the most cursed infra setups to make everyone suffer
6
19
525
24,741
Excellent results! This will be really powerful! I think you have confused the amount of tokens used for post training though. Their SFT data was only 1.8B tokens and trained for 4 epochs (8B tokens worth). Still very impressive!
1
4
341
Replying to @osoleve
> and load balancing is made up to scare children I'm very concerned about the state of your data-intensive applications
1
1
22
a model can never learn research taste because research taste is driven by spite
5
54
834
27,417
a model can never learn research taste because research taste is driven by spite
1
1
52
1,907
so ... how many papers reported these evals?
oh yes a shthon function
6
1
47
6,374
oh yes a shthon function
Looking at the data is so funny because in like 10 seconds you can find mistakes that people have just been ignoring for years. like this clearly bad regex replace in multiPL-e MBPP, replacing py with rs, resulting in this hilarious instruction to write a "rsthon" function
4
1
28
10,128
Looking at the data is so funny because in like 10 seconds you can find mistakes that people have just been ignoring for years. like this clearly bad regex replace in multiPL-e MBPP, replacing py with rs, resulting in this hilarious instruction to write a "rsthon" function
8
4
83
15,767
cursor employees be like
1
28
2,335
This was his first slack message btw
5
1,266
JUST IN: Anthropic CEO reportedly increasingly concerned that new hires are working there for a paycheck, as opposed to the company’s mission.
2
1
26
1,610
This is a good thing
This Yale + University of Chicago paper shows that real gap between LLM generated research ideas vs humans is not idea quality, but idea range: LLMs think narrower than human researchers. The researchers built a controlled test from 11,683 real papers, using each paper’s nearby prior work as the shared starting point. They asked models to propose a new motivation and method from those same prior papers, then compared those ideas with the real human paper ideas. Instead of asking whether 1 idea looked novel, they labeled each idea by what gap it noticed and what kind of contribution it made. Human ideas spread across many patterns, such as explaining mechanisms, testing failures, measuring evidence, building systems, and improving efficiency. Only 12.1% of human ideas were mainly about connecting separate work, but 47.1% to 64.2% of LLM ideas did that, meaning models used this move about 4 to 5 times more often. Even extra reasoning made this pattern stronger, suggesting models often polish a familiar recipe instead of finding more varied research moves. --- – arxiv. org/abs/2607.01233 Title: "Measuring the Gap Between Human and LLM Research Ideas"
3
1
27
4,884
Tell me, o muse, of a liquidated man
3
1
43
2,522
You have no idea the complaining we got when we made a radar plot for MPT 30B and the eval gauntlet. We just thought it was cool and reminded us of Japanese video games. About half the internet _did not_ like that
My favorite capabilities visualization is this one from the @thinkymachines Inkling release page. It's a good way to illustrate the spikiness of artificial intelligence and it looks like it gives away differences in priorities in training of different models. Quite good!
2
1
15
2,051
Finally got side tabs. So dope. It will be hard breaking a literal lifetime of muscle memory, though.
5
15
1,066
Wow, that is impressive. Tokenization is now free.
8
1,092
My managing agent presenting me with 3 PRs that its subagents constructed.
1
30
971
“We have j-space at home” J-space at home:
1
6
259
Sol 5.6 managing my other agents
14
731
Had the same agent session write up a report about the experiments it was managing. No one (but me and that agent) can read it
2
1
7
1,540