here for bad memes and overhyped tech.

MUC
DHH is a few weeks away from realizing Java is also great when you don’t look at the code
49
32
851
24,088
mx_twit retweeted
GPT-6 just cracked a 60-year-old fusion physics conjecture. For six decades, physicists have wrestled with the math governing magnetic confinement in fusion reactors. The equations describing plasma turbulence and heat loss are notoriously chaotic. A foundational conjecture proposed in the 1960s remained unproven. Human mathematicians hit a hard wall. The complexity was simply too massive to untangle. Until now. Researchers pointed a frontier AI model at the problem. Instead of blindly guessing, the model generated a rigorous, step-by-step mathematical proof that verified the 60-year-old hypothesis. It didn't just calculate numbers. It discovered abstract mathematical structure. Clean, limitless fusion energy is the holy grail of human civilization. The single biggest bottleneck has always been our inability to fully model and control high-temperature plasma. If AI can solve six-decade-old theoretical physics problems in its spare time, the timeline for commercial fusion just shifted forward. We spent the last century trying to force the universe to reveal its secrets through human intuition. Now, we just have to ask the model.
5
33
153
6,828
mx_twit retweeted
Math is easy now
6
6
72
7,958
mx_twit retweeted
Lattice Deduction Transformers has been accepted at @NeurIPSConf! 🎉 We're excited to present this work and grateful to everyone who has engaged with it so far. Here's a glimpse of LDT in action and how it works. We’re also preparing an improved revision of the paper, more soon!
Introducing Lattice Deduction Transformers: An 800k-parameter looped transformer that reasons like a SAT solver achieves 100% on Sudoku-Extreme with only 15 minutes of training. A collaboration between @axiommathai, @AmherstCollege and @BarnardCollege.
5
36
374
28,777
mx_twit retweeted
Die Energiewende kostet eine Familie soviel wie eine Eisdiele.
Ja, bei einer 4-Personen-Familie rund 232.000–260.000 € über 25 Jahre. Das liegt im Rahmen der typischen Gründungskosten einer Eisdiele (ca. 80.000–250.000 €).
9
175
804
17,567
mx_twit retweeted
Guys it’s over you don’t have to do this anymore
leetcoding at 33 years old is a humilation ritual
51
84
4,924
306,499
mx_twit retweeted
Singapore realising their education system raises NPCs with the survival skills of a baby panda and cannot survive outside a controlled environment. Solution: just plan their whole life as if it were secondary school. Next up, points for participating in adult CCAs and CIP.
I CALLED IT. It's happening y'all. The singapore gov has launched a pilot programme for a gov-run dating service. Looks like the pilot is currently restricted to just government workers for now.
52
216
2,090
177,077
mx_twit retweeted
Adam Neumann is a very problematic founder, but as a business thinker he is unironically excellent.
Adam Neumann on why top-down management fails with great talent: "Top to bottom means you think because someone reports in to you, they have to do what you say." "But it's not true. This is not a dictatorship. They have free will." "If you hired someone, they're really talented, they'll go work elsewhere. The more talented they are, the more the top to bottom does not work. They won't be willing to do it. They shouldn't." "Great talent is not willing to be managed like this." "If you're gonna manage these employees and you wanna attract the best talent in the world, remember that power comes from influence, not control." "If you think they need to do what you said because you're their boss, you've already lost, and it's a matter of time till it doesn't work out. If they're okay with it... they're not the right employee for you, and you're not the right boss for them. There's no chance you're getting the best out of them." @AdamNeumann w/ @StevenBartlett
24
51
1,174
565,808
mx_twit retweeted
The reason the US succeeds is because we have people with biology Ph.D.s who can quit on a dime and go work in software, not because we have them locked into aerospace engineering at 14.
A public high school in China’s Shenzhen that focuses on future aerospace engineers. The school has a quantum computing center for middle schoolers, plus AI labs, science labs, and future-engineer workshops… Special courses here include satellites, rockets, Mars rover…
279
478
11,742
1,204,850
mx_twit retweeted
the ending is always the same
32
92
1,853
51,728
mx_twit retweeted
CS academia is dead. So where does that leave AI PhDs like me? Did my ICLR reviewer bidding today. Skimmed a few abstracts, and the methods are the exact same recipe I learned when I got into 3DV two years ago. Swap in a newer video gen base model and boom, new paper 🥲 Makes you wonder how many of the 60k submissions were actually thought up by AI. Meanwhile, World Labs' Atlas has basically solved 4D scenes. Academia is so far behind it's not even funny. I still remember the day GPT-6 Astra dropped. My feed was flooded with GPT + Blender doing inverse graphics and GPT driving robot arms through manipulation tasks. The results were so good I literally had to sit down. A year ago, I was dead sure LLMs could never have spatial intelligence. A year later, Astra slapped that belief right out of me 💀 There's no doubt Astra was post-trained on tons of 3D and manipulation data, and it's only going to get bigger and faster. To me, that means any domain that can be represented symbolically, with clean benchmarks for RL, is going to get swallowed by LLMs. Next to real LLM intelligence, most academic papers that add a bit of inductive bias and tune their way to SOTA are just roadkill waiting to happen. Sadly, these papers keep piling up and flooding every conference. It's inertia, plain and simple. We've been chasing SOTA for so long that it's hard to stop overnight. But make no mistake: the paradigms in a lot of areas converged long ago. What used to be research is now pure engineering, and the room for academics to add inductive biases is only going to shrink. So as a PhD student, I see two paths left. One: if you can't beat them, join them. Clean data, build infra, then go to industry and train foundation models. Two: go back to real science. Stop caring about squeezing out another 0.1% on a benchmark, and start asking why this works and that doesn't, with foundation models themselves as the object of study. As a friend put it: AI research might end up looking more and more like biology, except the organisms are silicon-based 🤣 Beyond these two, it's hard to see anything that won't get eaten by LLMs. Still, I'm pretty pessimistic. I can feel the value of human knowledge being eroded bit by bit. Everyone will get their own AlphaGo moment. What will academia even look like after this... 😭
38
124
1,185
164,289
mx_twit retweeted
boomer origin stories are always like "Finally at 35 years old I decided to get serious with my life -- fortunately there was a national shortage of hedge fund analysts"
Replying to @skhetpal
He kind of got into investing by accident. He landed in merger arbitrage and special situations work, found out he loved it, and in 1988 joined Seth Klarman's famous Baupost Group. He spent a decade there learning deep value investing before striking out on his own.
76
755
17,066
1,410,102
mx_twit retweeted
PSA: Math and coding are the two most black-and-white fields there is, and AI's success in them does not easily generalize to others. Jumping from "AI solved a math problem" to "AI will kill us all" is not just absurd; it's deranged.
64
84
713
49,447
Replying to @Altimor
Doesn't work Never will Natural language by design isn't sufficiently precise
6
2
9
mx_twit retweeted
GPT3 writing code. A compiler from natural language to code. People don't understand — this will change absolutely everything. We're decoupling human horsepower from code production. The intellectual equivalent of the discovery of the engine. player.vimeo.com/video/42681…
138
609
4,618
Holy mother of EV transition
81
236
2,477
212,841
Black Forest Labs founder @robrombach: the startup is registered in the US. Why? A pure German GmbH wasn’t an option: Registering the GmbH took longer than training their first model, Flux 1. That’s just too slow. They need top investors – and for that, a US company is the gold standard. source: handelsblatt
37
104
991
124,201
mx_twit retweeted
Yesterday I asked Astra to categorize several business meal expenses and to add more description based on short snippets I provided. It did not add more description and instead spammed some slop into my spreadsheet, and I pointed this out, and it apologized. Do people with the view below actually do any real work whatsoever?
i basically think this is the Endgame. we've reached the point where another "capability doubling" over the course of the next four months brings us to very strongly superhuman performance in numerous areas things have *felt slow* for a long time because absolute capabilities have been well below many thresholds, but now that's just not true, and absolutely nothing indicates progress is slowing down. if anything the opposite. i can deeply hope that i'm wrong, that things stall out, that 2027 looks more like very impressive but still normal growth. but i just don't really believe that. or rather, i think the only way we get that is if we successfully enforce slowdown, which is itself a significant project. i think our default path looks like RSI by the end of 2027. this is the Endgame. the next *months* determine how it goes. if you have been waiting, if you have been worrying, if you have been thinking "maybe i should start figuring out how to help with ai safety", i think now is basically it. this is the last point where there's time to make an impact, the last point where u still can pull off a career shift and become effective and make a difference before it's finished. if you are in a phd program, or even undergrad, if you're working a job you're not too enthusiastic about that pays the bills, if you've been on the sidelines with capital or political power or in the labs not quite pushing as hard as you could this is the Endgame. now is the time. play all your cards.
72
36
1,119
140,574
mx_twit retweeted
Since there's a lot of RLM fandom going around, I hope no one gets mad at me for asking the obvious. I read the paper, I read the blog post, and I worked through some examples. I am still scratching my head about: 1) what's new here than what other agentic harnesses (CC mainly) are doing, 2) what's one straightforward example where RLM wins that all models and agentic harnesses fail? For example, the "quickstart" example feels like an RLM counterexample.
23
9
246
23,166
mx_twit retweeted
whenever something is crazy hype and is followed up by an armada of PR, you gotta bring out the bullshit detectors so i cloned the repo and did a little digging and found a lot of sound and fury. please enlighten me if any of these facts are untrue, would love to be educated RLM claim either overblown or false this doesn't even appear to be RLM (arxiv.org/abs/2512.24601)? RLM is "cool" as a concept because there's a premise of going infinitely deep to do arbitrary decomposition of a problem. in practice though it's not really hard, you just make subagents callable as functions and wrap that in a code execution tool. the reason RLMs have not been productionized and is just seen as a research ditty is that 1) going arbitrarily deep down the stack is a shitty thing to do in production without guardrails, 2) most problems do not need more than 2 layers of decomposition, and 3) agents are bad at banana phoning each other and will fail to preserve the subtleties of human requests, leading to a lot of inefficiencies and bad behavior. to solve these problems, there needs to be a fair bit of innovation, either on model training end or on harness end, to mitigate these issues. so what were the innovations in prime agent? they set RLM_MAX_DEPTH = 1? ... ok what? so theres no innovations, it's just calling subagents like every other harness in existence. if RLM MAX DEPTH is 1, it's not even RLM? it's just... a harness? so now i try to look up to what depth the ARC AGI 3 hillclimbing was done at. not surprisingly this doesnt seem to be disclosed on the blogs / posts. if it's just 1 layer deep, that's just shady marketing. a real innovation or contribution to harness engineering would be to detail the things that were improved to get a max depth of >3 to work without blowing up your computer, wasting a bajillion tokens, taking forever, or having agents go off the rails. i didnt find any 97% is just overfitting after hillclimbing on a public eval ok folks. it's nice that opus5 does better than sol than terra, but at the end of the day, we're looking at a public eval in which problems and solutions are open access, so any agent harness can arbitrarily overfit to the task set however much it wants to. if you do a codex or claude loop with some semi shady prompting to just get 100%, im pretty sure it will too. now obviously this is almost certainly not what the prime intellect team instructed their agents to do (otherwise you'd see that terra get 100% after a couple iterations), let's just all remember that there is no train / test split, and that ARC explicitly says - Public-set scores are vulnerable to task-specific overfitting. - They are “emphatically not” valid evidence of progress toward AGI. - Real generalization should be tested on the 55 semi-private or 55 fully private games. It's difficult to assess the degree of overfitting in Prime Agent's self-improving harness (which, btw, is not terribly different from nous agent or really any self-improving system that just looks at old JSONL agent trajectories and suggests skills / memories / system prompt appendments / tools / extensions). the more overfitting, the less impressive the 97% performance is. a counter-example that i found to be legitimately impressive was github.com/alexisfox7/PRO-LO… -- minimal / almost no overfitting, stupidly simple and general solution, and 95% on ARC AGI 3. i think the fairest thing to say about this system, given it's independent daemon system for managing subagent lifecycles, is that Prime Agent contributes a fairly robust persistent asynchronous agent-process tree. It does not demonstrate a solution to scalable deep recursive agency
Introducing Prime Agent: A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficient and expressive through programmatic tool calling, context as a variable, multi-agent messaging, and a self-modifiable harness state.
68
66
1,253
214,743