building bio informatics research pipelines on agentic ai rails at molecule.xyz | 🚴‍♂ | born at 334ppm | thegoodclimate.substack.com

RT @georgrestle: Dass die Linke jetzt quasi zur größten Gefahr für die Demokratie hochgejazzt wird, während eine in weiten Teilen faschisti…
471
stadolf retweeted
Claude can now help you build evaluations and hillclimb on them. In this article, we share guidance on eval design & skills that Claude Code can use to improve your applications. claude.dev/blog/automating-e…
155
522
8,090
1,595,410
stadolf retweeted
The PeptAI team has been selected for Track 2 of the @adaptyvbio x @AnthropicAI Protein Design Competition! The competition brings together researchers from around the world to use AI to design new potential drug candidates for diseases that affect millions of lives. Track 2 includes Claude Max 20x accounts for three months, plus $500 in @Modal compute credits for the team. Excited to join researchers exploring AI for protein design, with experimental testing and results shared openly on @Proteinbase so others can access the data and build on what we learn.
We’re partnering with @Anthropic to launch the biggest Protein Design Competition in the world, challenging people around the world to use AI to design new potential drug candidates for diseases that affect millions of lives. The competition will feature five challenges, each focused on a specific disease or biological mechanism. Compared to previous competitions, it will be a big step-up in complexity and scale to push the boundaries of AI-driven protein design. Together with Anthropic, we’re sponsoring over $1 million in experimental validation, making it possible to test more than 5,000 protein designs in our automated lab at no cost to participants. Anthropic is providing an additional $1 million in Claude credits. All experimental results will be published openly on @Proteinbase, including designs that didn’t work, so anyone can access the data and build on what we learn. The competition is open to everyone and free to enter. It will feature 3 tracks: - Track 1 is aimed at expert protein designers, with up to 20 teams to be selected. - Track 2 is targeting life science academics and industry researchers. - Track 3 is open to everyone from tech enthusiasts to high-school students. By combining Anthropic’s models with access to our automated lab, we want to make it possible for anyone with a laptop and an internet connection to join the global effort to advance human health with AI. A big thanks to @Modal for contributing compute for protein design and to @TwistBioscience for contributing the DNA for the experimental validation! Sign up link below -
5
8
48
13,062
stadolf retweeted
This year's record setting El Niño will impact many different parts of the world, with some experiencing severe drought and others flooding. In a new analysis at The Climate Brink, I explore what impacts happened during past events and which are most likely to occur this year. (link to the full piece below)
22
151
501
25,903
In 2024 I had to explain my team members that "Aerodrome" is indeed the best option for swaps and initial market liquidity on @base. 2 years later Aero now seriously takes on @Uniswap's pole position, crossing L2/chain borders transparently. aero.xyz/articles/aero-launc…
1
63
stadolf retweeted
People who don't read the code have no idea how much complexity agents dump into their codebase. I still keep agents on a short leash, but sometimes I give them a broader task and more freedom - eg, with personal dev tooling. That's when they go completely nuts with defensive code and remind me just how far into the stratosphere a complexity can get. They are literally heroin addicts without supervision. I've iterated on my "don't overcomplicate shit" rule a lot. Maybe it's me, but I still can't make them reliably stop doing this. My guess is that labs would rather bias toward defensive code than risk missing an edge case. Balancing that against complexity is still a human job.
81
17
411
33,919
stadolf retweeted
Boys and girls, nobody is thinking anymore. Entire enterprise departments have thousands of people who are spending their days interacting with the LLM, faking "being productive", consuming tokens because it's a highly regarded activity, shipping lines of code with Cursor or Amp or Claude or whatever (except their brain). Even slides are not made thinking anymore! The person that allegedly built the powerpoint deck you are reviewing right now cannot even answer a couple of intelligent questions on the very thing he's presenting to you. He looks at you stuttering some nonsense, hoping you'll let him off the hook because he made it so obvious Copilot wrote the slides (of course, you already knew that within the first 6 seconds of glancing at them). There is some sort of comical atmosphere to it all. And an absurd satisfaction in watching this entire situation unfolding in front of your eyes. There is also approximately ZERO chance this whole thing won't spectacularly backfire, and of course good old consulting firms will be there to pick up the pieces. I wonder at what point we started to believe we could get away with doing valuable things without even thinking?! How did this happen, wow.
75
97
839
58,536
stadolf retweeted
I’m slowly starting to hate AI code review. Massive amounts of nitpicking and an ungodly amount of corner-case defensive programming. After a few rounds the intent of the code is lost under mountains of extra code forced by the reviewer. I’m starting to be more agressive with implementer agents pushing back heavily for the reviewer agents to prove why their nitpick will make the code better, only for the reviewers to backtrack in frenzy. SO WHY DID YOU RAISE THIS IN THE FIRST PLACE? All this because I still require that „this code must be readable and easily understandable”
128
24
834
88,901
Das darf doch alles nicht wahr sein. Der Tankrabatt 3.0 kommt! Zum 1. Oktober schmeißt die Regierung zum dritten Mal in Folge Mrd. Euro an Steuergelder aus dem Fenster, um den Spritpreis um paar Cent pro Liter billiger zu machen. Haben Merz & Co. nichts aus ihren Fehlern gelernt? Dass ein wesentlicher Teil in die Taschen der Ölkonzerne geht und dass Bleifüße und SUV-Fahrer am meisten profitieren? Doch was mich am meisten aufregt ist, dass alle die kein Auto haben, mal wieder leer aus gehen. Schlimmer noch: Das Deutschlandticket wird ab 2027 natürlich wieder teurer, weil dem Bund dieses Ticket leider nicht mehr als 1,5 Mrd. pro Wert ist, was vermutlich genau der Summe des neuen Tankrabatts entspricht. Es ist mal wieder ein absolutes Armutszeugnis, eine Mrd. schwere homöopathische Beruhigungpille für Verbrenner-Fahrer, sonst nichts. Wie kann man 3x hintereinander den immer gleichen Murcks beschließen?
475
698
3,972
121,487
stadolf retweeted
Today we are announcing that S&P Global has entered an agreement to acquire OpenZeppelin. Onchain finance is growing from an emerging market into core financial infrastructure, and the standards and rails our team and community built are becoming the rails of global finance. OpenZeppelin smart contracts facilitated over $37 trillion in value transferred, with the vast majority of the largest DeFi protocols, blockchain networks, stablecoins and tokenized funds relying on them. With S&P Global, we expect to accelerate the impact of onchain finance, backed by more than a century of trust in global markets, benchmarks, and risk frameworks. To our clients and to all the users of OpenZeppelin open source tools: • OpenZeppelin Contracts and all our open source applications and tools remain open source, free, and publicly maintained on GitHub. Building open source standards stays a core priority. • Audits, engineering work, and ecosystem programs continue with the same team, brand, quality, and customer experience, with what will be the added benefit of S&P Global's research capacity, market data, and institutional reach. For the last decade, OpenZeppelin has set the security standard for onchain finance. Today begins a new chapter for that mission, together with one of the most trusted names in global markets. Read the full announcement: openzeppelin.com/news/spglob…
299
310
2,317
457,225
stadolf retweeted
Thrilled that #paper2agent is published in @nature today! Scientific knowledge is traditionally stored in passive papers. Paper2Agent transforms papers into virtual authors that answer questions, apply its methods and collaborate w/ other paper agents to make new discoveries đź§µ
54
443
2,183
164,961
We built recursive meta-intelligence, an AI that creates its own scientific instruments, turns them into persistent worlds that an agent ecology with hundreds of AIs inhabits, and uses those worlds to discover mechanistic principles in one of the hardest classes of physical problems: how complex hierarchical materials (nested structures of matter that give rise to new function through organization) evolve and fail. The AI reasons across enormous spaces of possible physical trajectories, where every rupture changes what can happen next, and compresses those histories into principles (which humans can understand and design with) - complex chains of causal events, highly nonlinear, and intricate. Scientific superintelligence is tangible here - machine-scale exploration opening cognitive channels into complexity that has been extremely difficult for humans to traverse directly. AI builds the spaces in which its next level of reasoning becomes possible; a representation becomes an instrument, the instrument becomes an executable world, and that world becomes the substrate for further intelligence. Intelligence then grows by constructing new spaces to think in. The task we explored started from a seemingly simple prompt to explore a biological material system - the AI then chose the representation, mechanics and experiments, built a fracture laboratory to push materials to their limit, tested hypotheses and generated scientific conclusions. The swarm explored a combinatorial universe in which architecture controls function and every rupture changes the future state of the material. The AI discovered a compact principle that defines how multiscale material architecture can program the evolution of failure. Material placement and geometric order determine how forces redistribute, whether damage cascades or remains distributed, and whether function survives substantial flaws. For this discovery to happen the AI had to reason across long path-dependent histories, simulate alternative futures and compress them into generative invariants (model-based causal reasoning, counterfactual simulation and temporal abstraction applied to an evolving physical world). It is incredible to witness this transition to a new form of intelligence and capability through scaling swarms. A lot of positive will come out of this because it expands the human epistemic horizon as AI can traverse thousands of possible histories and return mechanisms compact enough for us to understand, test and build from. Intelligence compounds through its artifacts! A few lessons we learned: ▶️ Learning and discovery are flows through spaces of possibility. Early work has shown how backpropagation flows through parameters, reinforcement learning through action and consequence, and autonomous swarms through representations, instruments and executable worlds, bringing it all together. Flows create structure; structure redirects future flows in the recursive instrument. ▶️ Nonlinear physics actually defines a larger principle, where high-dimensional dynamics generate stable invariants; invariants become effective variables; those variables become the substrate for a new level of cognition. ▶️ Recursion then becomes level creation - one possibility space compresses into a principle, and that principle opens a larger space above it.
AI for science is one of the greatest positive forces we have, and I cannot think of anything more human than to understand nature and to use the power to create new technologies that improve our lives, civilization and allow us to reach beyond.
Article

Recursive Meta-Intelligence

We built a recursive AI that creates its own scientific instruments, turns them into a world inhabited by a massive agent ecology, which then reasons across vast, nonlinear spaces of possible physical

178
468
2,936
590,219
I'm sad to report that the AA collab between 8130 and 8141 (Frames) broke down last week, and Base and Ethereum are now going separate ways to implement different AA standards. I want to share some reflections on this collab and on the future of the EVM. For a long time, the EVM has been a unifying force between L1 and L2s. Thanks to a standard account model (EOA) and a standard transaction type (EIP-1559), users have been able to enjoy their wallets working seamlessly across EVM chains. Similarly, a AA standard shared across L1 and L2s would ensure a consistent multi-chain UX for smart accounts, including post-quantum (PQ) accounts which we will eventually all use. As the crypto industry matures, however, L1 and L2s are starting to diverge in the values they provide and the use cases they target: - For the L1, it's all about CROPS -- censorship-and-capture resistance, open source, privacy, and security. In short, Ethereum L1 wants to be the most decentralized programmable settlement layer of the world, which is what makes it a good base layer for L2s in the first place. - For L2s, it's all about scaling, customization, and compliance -- things that commercial and enterprise use cases demand, and that the L1 does not provide. These diverging needs have pushed the shared layer -- the EVM -- to its limit, and AA proved to be the breaking point. While L1 and L2s both value core AA use cases such as gasless transactions and passkey wallets, they differ sharply in what these features must comply with: - For the L1, AA transactions must be uncensorable, private, and quantum-resistant, which call for a transaction type optimized for PQ signature aggregation and privacy protocols, and an account model that can be freely programmed and extended by developers without permissions from the chain. These needs lead to AA standards such as ERC-4337, EIP-7701, and now EIP-8141 aka Frame Transactions. - For L2s, AA transactions must work at high scale, and they must be legible such that the protocol can enforce clear rules about what kinds of accounts/transactions are permitted vs not. These needs lead to AA standards such as Tempo Transactions and now EIP-8130 by Base. With the AA collab, the authors of 8130 and 8141 tried to define a shared standard that can work for both L1 and L2s. While we identified a number of technical solutions, they all required one side or the other to compromise at least a little bit on their core goals. But ultimately, Ethereum wanted to be the best version of Ethereum, and Base wanted to be the best version of Base, and while both sides acknowledged the benefits of ecosystem interoperability, it was ultimately secondary to the need for each chain to achieve their core goals. So separate ways we went, putting the burden on wallets to deal with the fragmentation that ensues. Now, just because we ended up with fragmentation doesn't necessarily mean it was a bad outcome. If reducing fragmentation comes at the cost of homogenizing chains to the point that they fail to solve problems for the users they care about, that would not be a price worth paying. While I was initially sad that the collab did not come to a successful conclusion, I took solace in the fact that both Ethereum and Base are now free to innovate on AA to the maximal extent in accordance with their own visions, unshackled from the need to accommodate the other side. If they execute well, and if the wallet community can bridge over the fragmentation, we may well end up with the best possible UX for the end users. So where does that leave us -- the broader Ethereum community including Ethlabs -- if we want to continue pushing for a consistent UX across EVM chains? I see two paths forward: - We can establish a coordination mechanism that encompasses more stakeholders than ACD itself (where only L1 client devs have voting powers), to govern shared L1<>L2 resources such as the EVM. That way, L2s can participate in shaping the EVM, as opposed to having to accept whatever the ACD decides, or being forced to fork if they don't like the decision (such as in this case with Base). - We can accept that fragmentation of the EVM and wallet UX across L1 and L2s is inevitable due to their conflicting needs, and dedicate our resources to building wallets and applications that can abstract over the differences. Indeed, the collab was an exercise in the first path -- we invited Base, Arbitrum, and other stakeholders to directly influence how native AA shapes up for the L1. While it ultimately failed in this case, I feel a better outcome could've been achieved if we had established a dialog between both sides way earlier, as opposed to well after L1 core devs had rallied around Frames. On the other hand, I also learned from this exercise that some differences are unavoidable, and indeed it would be counterproductive to overly pursue interoperability at the cost of differentiation. In other words, we sometimes just gotta let the chains cook. For that reason, I've also become more bullish about the second path -- building wallets and applications that can speak the native transaction types of each chain, and hide the complexity from users through UX abstractions. This puts a lot of onus on the wallet/application developers of course, but on the bright side, it's also an opportunity for wallets/applications to stand out and differentiate, by competing to provide great UX across chains despite the underlying fragmentation. Ethereum is the art of staying together while remaining different. We must accept that chains will succeed by innovating, and innovations will naturally result in differences. On the other hand, we must never give up on dialogue when we can achieve interoperability without compromising core product goals. As Ethereum and crypto grow to eat the world, the push and pull between innovation and collaboration will only intensify, and it's up to all of us -- builders across L1 and L2s, applications and wallets -- to determine whether diverse innovations will split Ethereum apart, or make it thrive as one.
68
74
574
110,811
stadolf retweeted
There are two ways AI progress could go very badly and that we must avoid. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve people. To ensure that, we need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities. Second, we could end up in a world with too much concentration of power. If an extraordinarily powerful AI is used by one person or company to impress their worldview onto everyone else, the results could be extremely dystopian. Avoiding these two threats requires walking a narrow middle path; for example, one country could gain too much power. Another example is one lab ending up with too much power.
The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot. We welcome a federal framework that sets consistent safety requirements for frontier AI. But we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence. Consistent rules to manage frontier risk so that we can maximize the benefits are a good idea (and we are excited by ideas like independent auditors). Years ago, companies like ours developed things like Responsible Scaling Policies and Preparedness Frameworks. Those were good for that moment, and focused primarily on the deployment of completed models, not what happens during their development process. Today's shift to focusing on safe development and evaluation will need new tools. For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability, in addition to the safety work we have long done in advance of model releases. We hope that other companies will learn from our approaches and propose their own; we think shared standards for misalignment, monitoring, and safety will lead to better outcomes. We look forward to collaborating with our colleagues across the industry to formulate the best version of these. When we talk about “pacing”, we do not mean “stopping”. Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs. Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness, or let capabilities get ahead of alignment and monitoring. Where we will need the help of our government is for international coordination. But first we should do what we can ourselves.
4,719
2,183
21,413
4,705,842
Contextually available deep memory is still a barrier to cross in agentic operated systems, particularly across harness, machine, session boundaries. Funes seems to address that quite consequently: huggingface.co/blog/funes?ut…
56
Diese Woche war ich mangels Alternativen in einem Hotel direkt am Berliner Kurfürstendamm - eine Straße, die sich jeden Abend bis spät in die Nacht in eine Rennstrecke für getunte Protzautos verwandelt. Mit röhrendem Motor rasen diese Poser mit weit über den erlaubten 50 km/h durch die Stadt und terrorisieren mit ihren Lärm alle Anwohner der Gegend. An ruhigen Schlaf oder offene Fenster ist hier nicht zu denken. Wie kann sich eine Stadt, wie können sich die Menschen das täglich gefallen lassen? Wie darf es in einer zivilisierten Gesellschaft überhaupt sein, dass einzelne Männer ihren automobilen Fetisch auf Kosten so vieler Anderer ungestört ausleben dürfen? Warum gibt es hier kaum Blitzer, nächtliche Tempolimits oder temporäre Fahr- verbote, um den Anwohnern ein klein bisschen Ruhe zu geben? Ich hoffe so sehr, dass dieser Auto-Wahnsinn irgendwann ein Ende hat und dass die Politik, die all das auch noch als Freiheit verkauft, endlich abgewählt wird.
366
307
2,258
84,087
stadolf retweeted
MIT published a brutally honest report on what AI is doing to students. A committee of professors and students spent five months studying how AI changed learning on campus, and the findings read like a warning to every university on the planet. Study groups are disappearing. Office hours are emptying out. Problem sets and take-home exams no longer prove anything, because AI can produce credible solutions to almost any written assignment in the undergraduate curriculum. Students who lean on chatbots lose mastery and confidence, and some slip into what the report calls cognitive surrender, reaching for AI at the first hint of struggle. The numbers are rough. 46 percent of surveyed MIT undergrads use LLMs daily. 90 percent worry about their own overreliance. Undergrads who feel AI makes them replaceable now outnumber those who feel it makes them capable. The committee's answer surprised me. They refused to fight AI with surveillance. The report calls AI detectors unreliable, says lockdown browsers feel like spying, and warns that policing students builds a classroom atmosphere of mutual distrust. Instead, MIT wants to rebuild education around the things AI can't replace. That means oral exams, semester portfolios, in-person project work, and a required social component in every subject. The report even floats the idea of rethinking grades entirely, since without a GPA to optimize, much of the incentive to cheat with AI evaporates. The committee warns professors against replacing undergrad research assistants with AI agents just because they're cheaper, because a university exists to grow people, not output. The most famous tech school on earth admitted the machines broke its way of teaching. Its answer is more humans, not more software.
484
8,135
24,883
3,441,785
stadolf retweeted
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
10,657
16,396
87,873
76,558,504