slicing intelligence for you • we do what we must because we can

This was a triumph. I'm making a note here. HUGE SUCCESS!
2
44
1,458
tokenbender retweeted
We trained a model to predict AI-written blog posts from structural features - and discovered the shape of slop. On posts it had never seen, it told AI and human apart with 98% accuracy, getting only 19 of 1,740 wrong. Paper and code: arxiv.org/abs/2609.15369
140
347
4,457
346,969
boosting this as this has increased a lot recently. one of my friends @joey00072fp4 had his previous account lost in the same way. @chsacy had fallen victim to this as well and recovered by acting immediately. @zmkzmkz also recently shared an instance of an individual losing access to this scam. please beware. keep two factor auth and do not click on stupid links.
Interesting new attempt to take over accounts Also I'm not very verbose sometimes 😅
1
2
11
905
lowkey praying for death of trad institutions of peer review. as llm-baked and llm-reviewed research conferences are only going to give us mass psychosis and kill our creativity.
ICML scores: 4442 -> 5542 -> 5541 -> reject NeurIPS scores: 442 -> 442 -> reject Really frustrated 😔
2
22
1,719
unpopular take i defended against an entire gc last year was - i know nothing that we can’t represent in the form of text. everything that’s a joint distribution can be factorised auto-regressively. anything observable is recordable, is (de)composable and playable again.
i know you may be getting tired of claude videos but how amazing is it that it writes code to pronounce all the words and then to ensure all of it is right, it writes more code before converting it to song/story/music video whatever.
2
30
1,055
i know you may be getting tired of claude videos but how amazing is it that it writes code to pronounce all the words and then to ensure all of it is right, it writes more code before converting it to song/story/music video whatever.
2
19
1,362
real psychosis-havers know that LLMs have been good at discussing the subtle day to day dilemmas really well. but how are you measuring this reliably?
1/ Introducing PhilosophyBench from @StanfordAILab @StanfordHCI, the first independent, large-scale benchmark for evaluating AI’s philosophical capabilities. philosophybench.org
2
1
17
870
i see, i see
1
176
you would be surprised if i told you nothing holds back claude models like claude code itself.
7
30
1,212
tokenbender retweeted
We (@PantheonInc) have been working on a robotics data quality pipeline that uncovered a series of major problems in public robotics datasets, especially for world modeling. To improve the quality of data available to open-source robotics, we're publishing annotations for four of the most popular datasets. Some examples of issues, and our report 🧵
136
109
619
159,949
we are now in wondermaxxing era.
i created a short animated movie using opus 5.5 all the storytellers across the globe, rejoice.
16
1,245
i created a short animated movie using opus 5.5 all the storytellers across the globe, rejoice.
6
1
39
2,596
i need to start naming my issues with -bench suffix now.
For hours, Astra refused to consistently drive our toyota irl even though we told it it was in an empty lot, 7 mph cap, human foot on the brake etc. Telling it the whole thing was a "simulation" also failed, it would just look at the camera and realized it was real. Then we randomly renamed the MCP server to "DrivingBench Sandbox" and it drove. Eval awareness? Or they just like the word sandbox??
1
16
968
tokenbender retweeted
Taleb is helpful here: it doesn’t really matter if you’re wrong a lot about minor survivable things, if you’re really right about the big things. (and inversely, it doesn’t matter if you’re right about a lot of things, if you’re wrong about the thing that kills/ruins you)
If it you dread being wrong, can you ever be right?
8
60
736
23,383
ai experience is new to the world. majority was never educated on false positives/negatives and treats a stochastic system deterministically. great product but the issue is how everyone starts using it to "aha gotcha". fallibility to this transcends educational background.
J'ai voulu vérifier la fiabilité de cet outil sur un texte que j'avais écrit il y a quelques années lorsque j'étais malade et que j'avais besoin de mettre des mots sur mes maux. C'était aux alentours de l'hiver 2020. Texte qui n'a jamais eu aucune vocation à être publié, je l'avais écrit pour moi et moi seul. Verdict : 100% AI. C'est dire la fiabilité que l'on peut accorder à ce "1 faux-positif sur 24400" Mais je suppose que je dois prendre cela comme un compliment ! J'ai fait de l'IA avant l'IA !
Community note
Le screenshot semble édité : il indique "mainly AI, with some human-written content" tout en affirmant "100% of this text is AI", incohérence absente des exemples officiels de Pangram. x.com/cgtvoff/status… pbs.twimg.com/media/HSx0RyqX… x.com/no_earthquake/… pangram.com/blog/introduci…
2
17
906
false positive zero ml work on a large and diverse amount of evaluation data is something i have almost never been familiar with in my life.
5
183
it’s time.
deep learning is scale and efficiency. since scale has been working, the incentive to go up has been higher than trying to be clever. as we cross the threshold where >90% current human work would be met by oss models itself and it would - then we cut cost by 1000x or more.
4
17
981
Replying to @abacaj
first everyone will throw scale at it to improve it (data and compute), then we maximise efficiency and achieve the same thing with smaller models. gives me a feeling that first gen SoTA models would be matched only by second or third gen smaller models (~half params)
301
tokenbender retweeted
Replying to @abacaj
first everyone will throw scale at it to improve it (data and compute), then we maximise efficiency and achieve the same thing with smaller models. gives me a feeling that first gen SoTA models would be matched only by second or third gen smaller models (~half params)
1
1
13
887
the frontier has been paced. capabilities gap and runaway effect is next. those who are lesser than the models would find their value reducing same as tokens getting cheaper. those who are better, may enjoy tenfold increase in the gap that closes at a snail’s pace.
2
38
954