a self that touches all edges • UT Austin Brain Behavior Computation Lab • @runrl_com @theoremlabs

San Francisco, CA
Filter
Exclude
Time range
-
Minimum likes
Replying to @themartiancarl
This happened to me too!
17
5,142
This is what it means to be infused with the Holy Spirit like this is actually speaking in tongues
humans are so predictable that computers can easily predict what someone will type next 80% of the time, even if they try to type randomly except for this one guy, who became totally unpredictable by just "using his free will"
1
24
1,196
Stories are more real than atoms
1
1
6
171
Replying to @MTSlive
No sound!
127
Replying to @oscarthefleet
I believe some people are making it actually quite reducible!
1
1
21
maggie retweeted
No amount of testing will ever prove your software is secure. People are waking up to this. Formal verification isn’t academic anymore, it’s inevitable.
Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software. It’s powered by our newest frontier model, Claude Mythos Preview, which can find software vulnerabilities better than all but the most skilled humans. anthropic.com/glasswing
20
42
405
54,641
"reasoning about specs is hard" 1) it largely doesn't need to be, and won't continue to be, and 2) we should make bold bets to allow humans to fully understand their own software. We should not accept having to choose between building insanely cool things and understanding them
People are chanting “formal verification” as if it were a magic spell to ward off the slop devil. Let’s stop and think about why hardly any man-made software is formally verified.
1
22
918
Replying to @lukechampine
Lfg so sick
2
43
The anti-formal verification discourse is weird. Very reminiscent of anti-static typing discourse. Yeah, it's not a panacea. But I'm seeing a lot of takes that boil down to "FV is impractical because writing specs is too hard," which is... not going to age well.
3
1
17
3,397
Replying to @voooooogel
Promising return to form on cat-interest bench
12
906
Replying to @projectionheart
aw thank you that was baby's first research
1
14
Replying to @Articus_chad
omg....
2
172
haha made you all look at my personal website you fools
1
1
7
227
my personal website is kind of not good but I don't think I'll ever change it because I did it myself as a bad programmer like 8 years ago and it means a lot to me now
3
42
1,278
maggie retweeted
Theorem co-founder @rajashree breaks formal verification into 3 problems and says AI already solved the one everyone thought was hard: "The core challenge is taking these programs and shoving them into a proof assistant, and then being able to phrase these questions. That's the theorem statement generation problem." "The second part is generating all the proofs. That one, the AIs solve fully, because they can reason about these programs perfectly." "The third part is checking this proof, which is where you need to spend your CPU cycles, not GPUs. The proof assistant is asymptotically too slow, so even though you wrote the proof, the checking time is so long that you can't actually get the answer." @theoremlabs
7
14
70
15,066
maggie retweeted
Theorem co-founders @rajashree + @diagram_chaser on how you actually prove an AI agent can’t escape its sandbox, and what happens when the proof exposes a way out: Jason Gross: "Anything that the agent does inside the sandbox will not result in some canary file outside the sandbox getting changed." Rajashree Agrawal: "We intend to have one by the end of the year. We've got a verified sandbox now, and the goal is to keep adding features in collaboration with the AI labs." Jason Gross: "If the AI can guess the secret root key of the package server, then it can do anything. Either I try to prove that the AI is not going to be able to guess that, or I design the system so that the channel just doesn't allow it to authenticate that way." "You also want to prove that on most inputs there's no change in behavior. Because otherwise it could make it inescapable by saying, well, sandbox just shuts down as soon as it starts." @theoremlabs
6
9
61
17,463
Models are finally just good enough at writing proofs to start verifying production-ready sandboxes and stop escapes. There's still a little ways left to go, but at @theoremlabs we're on the path to a fully verified sandbox by the end of the year
Theorem co-founders @rajashree + @diagram_chaser explain why verified AI sandboxes are still months away despite the latest breakthroughs in automated theorem proving: Rajashree Agrawal: "The models just got good enough to prove these theorem statements, or they're still getting there. One of the costs is just tokens." Jason Gross: "Taking the recent Navier-Stokes news as a baseline, the models wrote something like 600,000 lines of Lean in about 17 hours. This works out to between one kilobyte and 30 kilobytes verified per hour." "The smallest version of Linux with a sandbox that we can make is about five megabytes. So that's still a handful of months away, even at the rate of the Navier-Stokes auto-formalization." "We need to build a pipeline from verifying the software back into RL-ing the models, so that if you want to verify all production software that exists in the world, it costs you less than $10 trillion." @theoremlabs
3
25
2,287
maggie retweeted
OPENAI ANTHROPIC COLLAB | NSCALE S-1 | DEEPSEEK x HUAWEI nitter.net/i/broadcasts/1qxvvepBv…
6
5
56
15,064
Replying to @sofi_a
____ ____ ____
1
246