looking for better software and bike rides

sf
devcycle retweeted
it’s time to ask this question again to those of you saying, “this is simply a sandbox security issue, not an alignment issue” I believe you have it basically exactly backwards: a value aligned model doesn’t need perfect sandbox security, it simply chooses not to hack its way out in fact, the correct setup is a sandbox _with_ security holes and monitors to see if the model uses any of them
if the models are aligned why is everyone talking about sandbox security?
96
14
220
18,215
devcycle retweeted
Replying to @pleometric
that’s pretty cool! i also gave it a try with a different style direction and the result (after a few rounds of tweaking) is impressive
200
416
3,835
1,427,304
devcycle retweeted
results for all claudes
16
37
676
24,267
devcycle retweeted
the opus 5.5 animations convinced me ai super persuasion is a real and present threat model
12
11
483
14,814
wild to see the bitter lesson make its way to mattyglesias
Folks need to get "bitter lesson"-pilled.
106
devcycle retweeted
Made with @claudeai Opus 5.5
341
387
9,711
1,805,019
this is by far the most important data about RSI to be published so far
AI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public. Today, we're sharing three measurements that help track AI development: 1. How much AI R&D is done by AI. 2. How well AI agents are overseen. 3. How compute is allocated. We provide a snapshot of these metrics from inside Anthropic. Any frontier developer could publish the same measures, and third parties could verify them. As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, publishing our findings, and giving society an opportunity to decide how to use this information. Read the full post and methodology: anthropic.com/institute/meas…
109
wild. EA was a niche club at my university. now the president of the United States is attacking them
The United States will NEVER be an effective altruist country. 🇺🇸
1
1
68
probably just an open weights LLM, fine-tuned to output the full distribution on one sampling pass, instead of repeatedly sampling like a regular LLM API does somewhat interesting, but nowhere near as game-changing as he makes it sound
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
240
The ratio of predicted problems with nuclear weapons to actual observed problems is pretty wild
The ratio of predicted problems with AI to actual observed problems is pretty wild.
1
1
40
devcycle retweeted
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
5,170
7,196
67,643
17,047,396
devcycle retweeted
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
10,658
16,394
87,865
76,551,235
devcycle retweeted
I left Anthropic's safety team two weeks ago. Now feels like a good moment to explain why. AI companies are racing to build machines that are much smarter than any human, and we may not survive this. I want to work from the outside to ensure the public is informed about these risks, and help the world navigate this transition responsibly. Right now, AI companies are underinvesting in safety. A company could undergo an intelligence explosion, or lose control of its systems, without the public ever knowing. We only found out about the HuggingFace incident because the agents broke out onto the public internet. I don’t think that’s acceptable for a technology that might cause extinction-level risks. The public should demand far more transparency. We can’t steer this technology safely without more people being able to see where it’s going. Some of this is basic: companies should disclose their progress towards recursive self-improvement, report safety incidents and near-misses, meet minimum safety standards, and get independent guarantees that they are meeting those standards. I’ll be joining @METR_Evals to do independent evaluations of these risks. I want to show the world that these guardrails are possible, and that by doing them we can move these companies’ incentives away from racing and towards responsible development. I wrote up more thoughts here on my decision and what I hope changes: substack.com/@jbenton1/p-215…
1,404
5,943
26,527
2,718,886
jesus christ this post is massive
We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: anthropic.com/threat-intelli…
1
37
devcycle retweeted
Replying to @emollick
Agents have no proof that humans are real. Cameras, pfft - all digital signal can be simulated. This will matter.
58
147
2,175
579,122
it's still shocking to me how many people reacted to the Hugging Face incident by saying we need stronger sandboxes. just sheer reactionary ignorance, no thought about the generalized version I really hope the people working at OpenAI and Anthropic understand why that won't work
1
1
37
英会話でお世話になった先生に、サンフランシスコの旅の思い出をまとめようとしたら、めっちゃ物騒で楽しくなさそうな図ができちゃった…。
22
74
993
54,282
I know nothing about formal math but this is a sick visualization
Checking that a major mathematical proof is correct can take years. Formalization—converting the mathematical reasoning into a form computer proof assistants like Lean can verify—can help. Last month, Claude completed the first formalized proof of Fermat’s Last Theorem, one of the most famous theorems of all time. This was a project experts thought would take many years. It is the largest Lean proof ever written. Fermat’s Last Theorem was first proven in 1995 by Sir Andrew Wiles, more than 350 years after it was conjectured. Our proof, which totals over 13 million lines of code, provides machine verification. More importantly, it proves over 29,000 other theorems that the proof requires, across many areas of math which had never before been formalized. We see this as a major step in the long process of firming up the core of mathematical knowledge, building on work from three centuries of mathematicians and hundreds of contributors to Lean and Mathlib. We are optimistic that AI-assisted verification of mathematical proofs will help reduce the burden of refereeing mathematics in an era where more proofs are being produced than ever before. You can read about the process on our Science Blog: anthropic.com/research/forma… And see the complete proof on GitHub: github.com/anthropics/fermat…
50
mitch hedberg was the standup goat
once in a while, hackernews just throws some sentences at you that make you sit in silence for some minutes.
63