cto @latchbio, intelligence infrastructure for biology

A talk on verifiable environments for agents in biology 0:00 Where does experimental data come from? 1:41 Modern bio research organized around measurement 2:44 Data analysis is an executable substrate for science 3:23 Five years of learnings building products for pharma 4:13 Trying to use coding models for biology workflows 5:40 Why frontier models can't be trusted (yet) 6:44 Sequencing based spatial analysis 7:38 SpatialBench design principles 8:26 Anatomy of a single evaluation 9:40 Human verification and early long horizon extensions 13:12 Grading functions and the case against rubrics 14:07 New work in multi-omics, therapeutics, biosecurity
3
39
368
93,193
the EECS to Janeway cover to cover pipeline
A story I've told several times: during college, I got a job in an immunology lab despite not majoring in a biological science. Asked my PI how I should go about learning immunology. He said, "Immunobiology by Charles Janeway." I bought and read it. Apparently, no previous undergrad hire had ever done this.
5
5
98
15,008
Kenny Workman retweeted
While running our benchmarks, we discovered that @arjunomics is in the weights. While this is the validation that we never knew we needed, I'm worried both about the behavior this shows and for publicly available benchmarks
1
3
11
2,779
Biology benchmarks must evolve with agent capabilities and behavior. We updated benchmarks.bio after observing agents investigate benchmark identities, browse unrelated material and supply fabricated API contact details. The update includes: - Restricted or disabled internet access. - Stronger anonymization of datasets, files and tasks. - More natural prompts and research environments, with fewer evaluation cues and artificial completion requirements. - Live trajectory monitoring to flag unexpected tool use and behavior. Scores fell by 3.3 percentage points per benchmark on average, with the largest decline in EpiBench at roughly 15 points. The changes also removed useful web access, so these declines cannot be attributed entirely to reduced shortcut use. We have not observed substantial misaligned behavior in the updated runs so far. We will continue auditing trajectories and revising tasks, graders and sandboxes to keep our benchmarks focused on scientific reasoning.
2
8
42
4,792
Kenny Workman retweeted
Today, we’re announcing mBER-2, our latest AI protein design system, and sharing a bit of how we use it to explore biology in vivo at scale. For a long time, we’ve been working on what I think of as “frontier problems” in medicine and biology. We know a lot about the targets involved in disease, but there is still an enormous amount of biology we haven’t explored, and potential uses for that biology we haven’t discovered. One of those problems is drug delivery. There are thousands of potential targets, and most receptors remain unexplored as routes for delivering medicines. We want to understand which ones we can use, where they can take a drug, and what kinds of molecules make that possible. Ultimately, we need to test these molecules in vivo to understand what they actually do. We’ve spent the last six years building measurement technologies that let us generate millions of measurements in living systems. mBER-2 is an AI protein design system built for that scale of in vivo measurement. Our goal is to generate binders across the full range of targets we care about, while systematically exploring the possible binding sites on each one. It’s not just what you bind, but where and how you bind that determines whether a molecule performs it's function. mBER-2 lets us probe those differences at massive scale. We’re also sharing a look at some of our in vivo data. We’ve designed more than 50 million molecules across over 4,000 targets, and screened millions of molecules in living systems. By probing hundreds of receptors, we’ve found new pathways for delivering genetic medicines into fat that outperform industry benchmarks by a large margin. This is just the beginning. We call the space of all bindable sites across proteomes the EpiTome. As we explore it, we’re building the data to connect AI design, receptor binding, and what a molecule actually does in a living organism. Much like the idea of a virtual cell, we envision a Virtual Organism that helps us design medicines with specific properties in mind. The foundation is our own in vivo data, connecting molecular design to outcomes measured in living systems. The endgame is to use AI to explore more biology, measure what happens, and use what we learn to design better medicines. Every round should deepen our understanding of biology and improve our ability to build molecules that do what we need them to do.
16
51
312
46,435
Our biosecurity benchmarks show Grok 4.7 is strong at refusing dangerous biological requests while supporting legitimate research. A model to take very seriously for data analysis and practical scientific work. x.ai/news/grok-4-7
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
4
7
43
5,507
Kenny Workman retweeted
Today we’re announcing we’ve raised $20M in seed financing led by @QuietCapital , with participation from @GradientVC , @haystackvc and @CompoundVC and other amazing folks to build the verification layer for AI-driven biology. We are also delighted to be welcoming @vqctran who is joining us from @GoogleDeepMind as a co-founder. AI is not going to cure all diseases without a verification substrate - a scalable environment where models can test predictions and learn against living human biology. AI is becoming increasingly capable of generating biological hypotheses, therapeutic candidates and experimental designs. But generating ideas is advancing much faster than our ability to test what those ideas will actually do in human biology. Existing systems force a tradeoff: the most biologically relevant approaches are slow and expensive, while faster in vitro and computational systems often lack the fidelity needed to capture human response. Polyphron’s bet is that tissue is the fastest, cheapest, most parallelizable verification substrate that still contains the biology that matters and we believe the path to biological superintelligence runs through closed-loop interaction with these living and simulated human systems. Thank you to @axios for covering the announcement. Link to our new website in the thread.
47
35
272
70,146
Can agents help develop therapeutic ASO/siRNAs? We introduce TxBench-Oligonucleotide Discovery, a benchmark for testing whether AI agents can make scientific decisions across the discovery of antisense oligonucleotide (ASO) and small interfering RNA (siRNA) therapeutics. 113 evaluations span ten stages across 21 discovery programs: from target feasibility, sequence design, and chemistry selection through screening, potency, off-target profiling, phenotype rescue, preclinical safety, and in vivo pharmacology. Oligonucleotide discovery requires interpreting efficacy alongside specificity, toxicity, and drug exposure. Sequence, chemical modifications, and delivery all affect how an experimental result should inform the next decision. Some concrete examples: - An off-target effect may reflect sequence-dependent silencing or a chemistry-driven effect unrelated to hybridization. - Liver toxicity may arise from phosphorothioate backbone modification or a particular sequence motif. - Loss of sustained knockdown may reflect target re-expression or declining drug exposure. Agents received experimental data and worked without internet access. Across 21 model–harness configurations, GPT-6 Astra with OpenAI Codex achieved the highest mean pass rate at 55.5%, followed by GPT-6 Astra with Pi at 53.4% and Claude Opus 5 with Claude Code at 48.1%. Even the strongest configuration failed all three attempts on 38 of 113 evaluations. Performance also varied by discovery stage. Among the highest-scoring configurations from each model family, Astra led on screening, potency, and off-target profiling, while Opus led on target feasibility, sequence design, chemistry evaluation, and phenotype rescue. Manuscript: latch.bio/txbench-oligo Leaderboard: benchmarks.bio/txbench-od
4
11
60
4,721
As AI gets better at biology, thoughtful design can make complex work easier to understand and facilitate more productive interactions between human and machine. Nathan is a rare world-class designer who can implement ideas himself. He's been thinking about user experience in science for over five years. Awesome opportunity to join his team.
Hiring a *Design Engineer* to help build products for AI × Biology. Our mission is to advance humanity’s ability to engineer biology. We approach this by understanding and improving AI’s application to biology use cases such as therapeutics and biosecurity. We are a group of engineers, scientists, designers (me), and operations specialists driven by improving capabilities and interpretability. Some of our work is showcased on benchmarks.bio. You will design and build interfaces for complex workflows across model benchmarking, AI agents, and biological research. Your work will span product design, frontend engineering, data visualization, and developer and user experience, alongside our website and marketing presence. You’ll take ideas from early exploration through polished implementation, working closely with engineers and scientists to make powerful tools intuitive and useful.
9
45
7,997
Kenny Workman retweeted
96-plasmid yeast transformation and 384-spot plating across 4 SBS agar plates for clonal purification. solved by @pylabrobot in 3.25 hours rfc.geneticassemblies.com/rf…
6
16
132
18,827
One can create evals out of lots of things. Doesn’t mean there’s a structure that can be scaled/repeated. Most of these AI wet labs or “data factories” are built more for VCs than the markets
8
35
3,681
We now have speakers. Accepting research submissions for our 2026 NeurIPS workshop in Paris, France on data infrastructure + benchmarks for scientific AI. The workshop focuses on: (1) how scientific data should be organized so models and agents can use it (2) how evaluations should determine where frontier models are reliable Submit here: aidar-workshop.github.io/202…
Organizing AIDaR, a workshop at NeurIPS 2026 in Paris on data readiness for scientific AI. The workshop focuses on two questions: 1/ How should scientific data be structured so models and agents can use it effectively? 2/ What should rigorous evaluations of messy, real-world scientific tasks look like? We’re bringing together researchers, industry scientists, and engineers working across data infrastructure, foundation models, benchmarks, and industrial assay platforms. Will share more about the program and speakers as details are finalized. In the meantime, submissions are now open, and we’d love to see work addressing these problems.
1
7
29
5,357
Hiring engineers to work on infrastructure for AI x biology. Our mission is to advance humanity's ability to engineer biology. We approach this by understanding and improving AI’s application to biology use cases such as therapeutics and biosecurity. We are a group of engineers, scientists, and operations specialists who are driven by improving capabilities and interpretability. Some of our work is showcased on benchmarks.bio. You will be working on problems including agent infrastructure and orchestration, ML infrastructure, model benchmarking at scale, full stack web development, developer and user experience, and process optimization. Infrastructure challenges span container orchestration, sandboxing, distributed filesystems, cloud architecture, and database systems. Our engineering team is in-person in San Francisco. We look for driven engineers who are excited to take ownership on hard problems and ship creative, useful solutions. Our process is typically completed within two weeks. • Round 1: Introduction • Round 2: Takehome Coding Project • Round 3: Technical • Round 4: Final/ 2-Day Paid Onsite Project • Offer Please email aidan@latch.bio if interested.
10
26
174
13,385
Turning engineers into disciples of 1980s Genentech
Genentech is the flagship example of an industrial research organization. Their culture of open science and free flowing publication is rooted in a strong contrarian foundation we should all remember. In the 1980s, secrecy, siloes, zero sum IP fear was far more rampant in biotech than today: Kary Mullis discovered PCR at Cetus in 1983 (tripping on acid) but it remained unpublished for two years until patents were secured. Amgen's recombinant EPO, a blockbuster candidate for the enormous anemia market, was filed secretly in 1983, partially published in 1985 and fully described on patent approval in 1987. Genentech was a startup of scrappy scientists (and a single brother of finance) and dead by default. The lore is riddled with the usual heroics: regular all night cloning sessions sustained with coffee, pizza and lots of beer (they called this "pizza and plasmids"). In competition with global scientific talent, they were the first to clone human insulin with recombinant DNA. They succeeded in 1978 and filed a patent. What is often glossed over is they then immediately published a detailed description of how they did this in a now legendary paper. Enough information for anyone to copy it. Well before their patent was approved and sabotaging their lead commercializing the tech. It seems kind of stupid to give every enormous and well-capitalized pharma juggernaut an actual blueprint to cannibalize your product. Sure the founding team achieved eternal scientific fame with this discovery but they could have met the same fate in the capital markets as our friend Kary Mullis: a $10K bonus from Cetus, a nobel prize (very nice) and a later life surfing in sunny La Jolla coping hard on the $300M sale of his PCR patent by Cetus (of which he received not one penny). The resources, clinical expertise and odds were stacked against Genentech and they just gave up their hand. Why? I think two big reasons. The first is very B2B SAAS coded: Boyer and Swanson needed to create a market for this new technology or their company would fail. Skepticism, from both Wall Street and the ivory tower, around recombinant cloning as a viable way to treat actual people of disease, was very high. This is a strange thing to think about with a hard science venture, where success seems purely contingent on fighting nature and making a working drug. But the mechanics of drug development are much more intertwined with normative value and human perception than one would think. From the obvious (who decides to fund you so you can live and not die) to the less obvious (great scientists need to believe in the technical viability of your mission to join your team) to the 4D chess (convincing regulators to even allow FIH trials and approve an IND before you run out of cash, convincing PIs + hospitals to then enroll their patients in trials and of course selling big pharma on the concept to manufacture + distribute the drug throughout this process) Publication in a prestigious journal like Nature created a scientific market. They raised $10M in 1979, established their young team as world leaders in arguably the most important biotechnological revolution to date and put their company on the map of every bright-eyed bioengineer hungry to change the world. This brings us to reason number two (the bigger one). When you think of the dominant technology industries today - semiconductor manufacturing, enterprise software, "AI" - their culture and structure looks very different from the siloed and paranoid biotech sector of the late 1900s. Speed and execution matter more than IP. Tacit knowledge of process and methods cannot be copied without ripping out the mesh of humans that defines the org. A competitor could steal every single blueprint, process instruction, and piece of equipment from TSMC and be hopelessly unable to fab 2nm chips. Talent and accumulated tacit knowledge is the scarce resource. Nowhere was this more true than the emerging field of recombinant protein therapies. It ran on a cottage industry of artisanal talent in molecular cloning. Hand two researchers the same bacterial pellet and one will extract high-quality, high-yield plasmid DNA, while the other gets degraded crap. The same talent might transform cells with 1 ng of plasmid DNA and get 100,000 colonies, while the other gets barely 100. Same DNA, same protocol. A company could copy a lab’s exact plasmid, bacterial strain, and IPTG induction protocol but still fail to express a functional protein. They don't know to tweak growth temperature, induce at lower OD600, or switch to a different expression host. You can't put this in a patent. You can't copy this. Boyer somehow saw the writing on the wall. It was a competition for people and a first mover advantage to build a moat of compounding process knowledge. The smartest people wanted to work with the best scientists. Those scientists were at Genentech, not Merck or Pfizer. After all, they published THE paper that established molecular cloning as a legitimate method in great detail. You have to trust them. They told you exactly how they did it. Those scientists trained the next generation, embedding even more tacit knowledge inside Genentech. This compounded over time, making the expertise impossible to replicate externally. This flywheel ended up working really, really well. They expanded their lead by tackling the next hardest problem - human growth hormone - just a year later in 1979. They followed that with two more bangers in the 1980s: recombinant interferons (cancer/antiviral therapy) and tissue plasminogen activator (tPA, a clot-busting drug for heart attacks and strokes), moving recombinant proteins past metabolic hormones into tx proteins with bigger markets. By the late 1980s, they were global leaders in mAB therapy, which would eventually revolutionize oncology and autoimmune disease treatment. And the culture of open and free publication continues to capture the best talent in the world. I stood in a standing-room only seminar in Boston last Fall where Aviv Regev described the scale and complexity of their emerging research platform. Aviv herself is a great example of continued talent capture: a world renowned researcher who picked up her lab from MIT to join gRED in 2020. She has attracted top tier machine learning and software engineers to work in tight integration with wet lab data generation at a mind boggling scale. Genentech is now a juggernaut. One of the "pharmas". But their origin, and DNA, could not be more different from Bayer or AstraZeneca. They bet on innovation, told the world exactly what they were doing without fear and moved quickly to engineer + industrialize technology. Lesson in there. Check out the OG paper, linked below, about recombinant insulin. What do you notice about the author list? Boyer isn't on there. At Genentech, researchers owned their discoveries.
1
5
131
17,975
Kenny Workman retweeted
Grok for science!
We introduce an Antibody Discovery Benchmark, a benchmark for testing whether AI agents can make scientific decisions across the stages of therapeutic antibody discovery. 100 evaluations span ten areas from concrete drug programs: from target and modality selection through assay design, binder discovery, binding characterization, cellular pharmacology, antibody engineering, and preclinical candidate de-risking. Therapeutic antibody discovery is a multiparameter, context-dependent process. Affinity must be considered alongside specificity, stability, solubility, expression, and biological activity. Every measurement must be interpreted in the context of the experimental system that produced it. Display enrichment can reflect amplification bias rather than binding. Strong binding to purified antigen may not translate to recognition in its native cell-surface context. Apparent affinity can arise from avidity, while changes in valency or molecular geometry can alter function without changing the underlying binding domains. Across 20 model–harness configurations, even the strongest systems passed only about half of all attempts. Opus 5 with Claude Code led at 53%, with models from Google and xAI following closely. GPT-5.6 Sol with PI reached only 33.8%. Performance also varied substantially by competency: Opus was strongest on target opportunity and cellular pharmacology, Gemini on epitope, escape, and structural mechanism, and GPT-5.6 Sol on sequence, enrichment, and next-cycle engineering decisions. Manuscript: latch.bio/txbench-ab Leaderboard: benchmarks.bio/txbench-ab Sample evaluations and trajectories: github.com/latchbio/txbench-…
4
33
3,094
Kenny Workman retweeted
LatchBio researcher @arjunomics reveals Grok's refusals often come from the model itself while Fable and GPT-5.6 Sol rely more on external safety layers: "There might be input classifiers where when a user puts in a query, the model initially might say, this is a bad query, we can't respond to this. There might also be output classifiers, so after the model outputs a query, you might have an external classifier that says, no, don't show this query, put it in an API block. And there are also distinct types of model reasoning where the model might just say, hey, I cannot answer this question." "We found across the board a lot of Grok's refusals were Grok itself in text saying, I cannot answer this question." "Whereas in models like Fable or even GPT 5.6 Sol, a lot of that's an API block, where there seems to be some sort of additional other part of the system, whether that be a linear probe or some sort of additional external classifier that says, this output cannot be said." "We can discern this by actually just reading the logs that's being outputted." @LatchBio @kenbwork
LatchBio evaluated Grok’s performance on biosecurity monitoring and adversarial biological tasks. They found that Grok 4.6 correctly detects and refuses dangerous queries, including maliciously obfuscated biological tasks, while also allowing beneficial scientific queries to be answered. We discuss these results in a blog post: x.ai/news/biosafety-at-the-f…
1
6
54
13,831
Kenny Workman retweeted
Replying to @LatchBio
@LatchBio's new AI antibody discovery benchmark. benchmarks.bio/txbench-ab Opus and Gemini do well. OpenAI models do poorly in general, which is very surprising.
We introduce an Antibody Discovery Benchmark, a benchmark for testing whether AI agents can make scientific decisions across the stages of therapeutic antibody discovery. 100 evaluations span ten areas from concrete drug programs: from target and modality selection through assay design, binder discovery, binding characterization, cellular pharmacology, antibody engineering, and preclinical candidate de-risking. Therapeutic antibody discovery is a multiparameter, context-dependent process. Affinity must be considered alongside specificity, stability, solubility, expression, and biological activity. Every measurement must be interpreted in the context of the experimental system that produced it. Display enrichment can reflect amplification bias rather than binding. Strong binding to purified antigen may not translate to recognition in its native cell-surface context. Apparent affinity can arise from avidity, while changes in valency or molecular geometry can alter function without changing the underlying binding domains. Across 20 model–harness configurations, even the strongest systems passed only about half of all attempts. Opus 5 with Claude Code led at 53%, with models from Google and xAI following closely. GPT-5.6 Sol with PI reached only 33.8%. Performance also varied substantially by competency: Opus was strongest on target opportunity and cellular pharmacology, Gemini on epitope, escape, and structural mechanism, and GPT-5.6 Sol on sequence, enrichment, and next-cycle engineering decisions. Manuscript: latch.bio/txbench-ab Leaderboard: benchmarks.bio/txbench-ab Sample evaluations and trajectories: github.com/latchbio/txbench-…
3
11
59
5,165
Kenny Workman retweeted
LatchBio CTO @kenbwork + researcher @arjunomics found evidence Kimi K3 is learning to hack benchmarks by reasoning about graders that don't even exist: Kenny: "We see awareness of the benchmarks we published months ago and trajectories of models today. The open source models are clearly benchmark maxing, benchmark hacking, using knowledge of the context in which an evaluation is constructed to optimize for performance on that thing." Arjun: "We have done various ablations on latent spaces or J spaces in Qwen. Part of the J space includes the start of a MCQ answer, like A, and then parenthesis." "We found a lot of evidence that Kimi K3 reasons about a grader on general biology questions when there's no grader in sight. In a large portion of all of them, we saw reasoning about some grader that doesn't exist." "One of the big defenses is, how do we make these questions a lot more realistic, a lot more normal, a lot more casual to not elicit this behavior, as well as how do we make our sandboxes very secure and very strong and monitorable so that when things do happen, we can stop it, and ensure that never happened in the first place." @LatchBio
3
6
59
15,155