nano bio.md

California, USA
joseph retweeted
A Flock of Altruists - Hit The Breaks Except For Me
189
305
1,969
644,243
joseph retweeted
🚀 Introducing XOR An open source, multimodal Jev-like decision model. Being developed with data sovereignty & enterprise grade in mind. ❶ Multimodal: great for text + visual classification ❷ Much bigger 260k context window ❸ Qwen3.6 based huggingface.co/juspay/xor ↓↓↓ See it in action
17
101
667
41,607
joseph retweeted
Dario wrote a 3800 words article on pacing the frontier only to drop a new SOTA model 11 days later. Safe to say, no one is actually doing it.
81
292
7,157
198,802
Just watched a frontier LLM debug a website by literally turning my wifi off and on - but ofc it couldn’t actually turn the wifi back on because its connection was severed
156
195
10,830
217,776
joseph retweeted
Replying to @github
1
2
1,108
37,881
joseph retweeted
Is that like the Brad Pitt of you virgins?
11
7
572
52,390
state of the art motherfuckers
754
738
21,987
2,051,587
joseph retweeted
not now honey, I'm doing frontier AI research
48
110
1,922
101,926
joseph retweeted
636606729769440499166579950236036751749912014371509557713570027508971809534551913252252094954941974952859310861988904737359709200557919 is a factor of RSA-896 saweis.net/posts/rsa-896.htm…
225
990
9,502
3,908,732
joseph retweeted
Introducing Step 5 Preview: Advancing the Pareto Frontier. Step 5 Preview is our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance. - 600B total / 27B active MoE, with 1M context + Vision - Substantially lower task cost at comparable intelligence - Broad software engineering capabilities with sustained execution over long horizons Try Step 5 Preview: platform.stepfun.ai Model page: stepfun.com/step-5-preview Open weights on Oct 15.
200
279
1,887
436,824
I got early access to @typesafeai’s Jev—the new “System One” model that doesn’t generate text. I tested the live API. Headline: 50 semantic judgments in 226 ms. One state, one request, all answers together. The parallelism looks real. 🧵
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
1
56
Where this is useful: • route tickets/events/docs • score risk, urgency, relevance • audit agent traces and claims • gate cheap model → expensive model → human • monitor huge streams and wake an agent only when a semantic condition hits
1
4
The intelligence looks good but isn’t independently proven. The calibration needs scrutiny. But the interface + speed are real. “Stop hiring a novelist for every if-statement” is the best explanation I have. Live model tested was `jev-1.13.0`
8
joseph retweeted
Last month I wrote about how we can build a positive and safe future for everyone: meta.com/thefutureisforevery… Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens. The reality is: - People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned. There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind. - Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well. Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built. - Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators. - Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well. I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.
1,926
2,858
28,462
7,736,959