codex has been super buggy today for example there's something up with the top left corner where it looks pretty ugly now
14
97
9,060
it's so beautiful...
15
29
385
21,245
kinda wild that we got the first good definition of AGI in the year 2000 and nobody actually tried training for it until 2026
2
1
52
3,439
it also seems like this should scale and be practical if you just pick any sensible coding language that can be used to access the internet fairly easily this is the coolest paper in a long time imo every greenfield continual learning neolab should be taking note
2
2
52
3,234
average OPSD replication effort: "it doesn't really work but wouldn't it be sick if it did?"
11
5
210
12,670
if i had to steelman the general usage of OPSD, my recommendation would be: - use judges to identify behaviors in rollouts which expressly violate guidance which is *already in context* (e.g. system prompt rules, tool schemas) - insert a verbatim reminder of that guidance right before the failure occurs - train only on the tokens immediately following the reminder this avoids hint leakage, and combats the legitimate problem of rule adherence in long-context settings. however, you can also just do this with RL, and it's not clear why you should expect OPSD to be any better than folding those same failures into reward penalties. it's also training the model to take actions which might directly conflict with preceding reasoning, which could be an issue for CoT faithfulness if that's something you're concerned about. tool schema failures are the clear win, but if your model is regularly messing up tool calls, there are probably more bigger problems in your post-training to solve. i think it's pretty unlikely that any serious frontier lab uses it in a real way other than as a last-minute band-aid on weird formatting bugs which show up in final testing states. it's not bitter lesson pilled. some people like to think that *specialization* isn't bitter lesson pilled, and i very much disagree here. the most intelligent and successful humans are usually exceptionally specialized and spiky. somehow we convinced ourselves that being superhuman at *everything* ought to be a free lunch, but there's not really any historical precedent or theoretical argument for this being true in the limit. the success of frontier foundation models isn't really evidence here, they're *way* more expensive to train than specialized counterparts, but final-run training costs are dwarfed by research and inference, and it's much easier to ship a one-size-fits-all product. jev is a hit because nobody ever made BERT finetuning easy enough for non-expert developers. real continual learning has never been tried. we don't even have the cognitive core nailed yet.
2
1
45
1,832
will brown retweeted
Thank you for the shoutout @MTSlive 😭 Here are some of the companies I'm the most excited about, need to add more! Big fan of someone finally taking on the immovable (and awful) wall that is LinkedIn Go @cosign @MTSlive @eriktorenberg @david__booth !!!
1
6
39
6,827
lots of categories emerging around the AGI stack: - open model RLaaS - data & evals - inference - coding harnesses - GPUs - agent tracing - sandboxes at @primeintellect, we agree. we do all of these things. people used to ask why we do so many things. they ask that less now.
11
23
463
17,651
what did you think open superintelligence stack meant? vibes? essays? we’ve had this as the north star for a long time, we’ve come a long way, and there’s so much more to build and research we’re hiring btw primeintellect.ai/careers
2
26
1,959
just realized a crazy method that’s not even close to possible rn but will be like next year damn
13
174
15,567
in 2024 i stopped writing papers and started writing tweets and that was such a good call honestly
21
9
599
35,590
i think this is my favorite piece of ai art ever it goes hard, it's heartwrenching and hilarious, and like all of the best pop punk, it's derivative and asinine and kinda sounds like shit truly the platonic ideal of "pop punk song made by a computer"
opus 5.5 just dropped its first pop punk single with a music video! everything you see and hear is generated from javascript code that claude wrote. no samples, no libraries 🔊
11
4
213
23,045
it reminds me of across the sea by weezer
8
1,019
i love how this is like 7% @primeintellect technical staff
the big labs and new media are black holes for talent and epistemic collapse here's to the folks on the outside
8
5
278
18,926
Muse is on track to be the second most universally beloved Meta AI product behind Instagram Ads
13
4
290
12,421
alexandr wang is the next jony ive
25
14
394
27,956
will brown retweeted
Highest throughput GLM 5.3 🫡 docs.primeintellect.ai/infer…
13
15
214
21,750
easy way to convince yourself this is true is to look at how many "notable benchmarks" from 12mo+ ago saturate at <90% the frontier models of today should get literally 100% on these kinds things...
it's becoming clearer that all models of the past couple years were trained on lots of bullshit broken data but managed to get good enough recently that you could point them at the data with careful orchestration and review loops and fix most of the issues and now it's kinda fine
19
1
294
28,359
it's becoming clearer that all models of the past couple years were trained on lots of bullshit broken data but managed to get good enough recently that you could point them at the data with careful orchestration and review loops and fix most of the issues and now it's kinda fine
28
23
812
56,638
one theory of mine re: new opus is that anthropic finally managed to properly scale "taste RL" which is noisier and more subtle than normal RLVR and just needs a crazy amount of rollouts, hence doing it on a "smaller" model + after getting more compute + inference gains
6
1
170
6,068