Intelligence Infrastructure for Biology

San Francisco, CA
LatchBio retweeted
benchmarks (dot) bio v1 has many improvements. It now includes an overall leaderboard (1105 tasks), counts refusals as failures, and makes patches to handle misalignments. It has also been re-tested back to Sonnet 4.6 (feb 2026).
Biology benchmarks must evolve with agent capabilities and behavior. We updated benchmarks.bio after observing agents investigate benchmark identities, browse unrelated material and supply fabricated API contact details. The update includes: - Restricted or disabled internet access. - Stronger anonymization of datasets, files and tasks. - More natural prompts and research environments, with fewer evaluation cues and artificial completion requirements. - Live trajectory monitoring to flag unexpected tool use and behavior. Scores fell by 3.3 percentage points per benchmark on average, with the largest decline in EpiBench at roughly 15 points. The changes also removed useful web access, so these declines cannot be attributed entirely to reduced shortcut use. We have not observed substantial misaligned behavior in the updated runs so far. We will continue auditing trajectories and revising tasks, graders and sandboxes to keep our benchmarks focused on scientific reasoning.
1
14
962
LatchBio retweeted
Biology benchmarks must evolve with agent capabilities and behavior. We updated benchmarks.bio after observing agents investigate benchmark identities, browse unrelated material and supply fabricated API contact details. The update includes: - Restricted or disabled internet access. - Stronger anonymization of datasets, files and tasks. - More natural prompts and research environments, with fewer evaluation cues and artificial completion requirements. - Live trajectory monitoring to flag unexpected tool use and behavior. Scores fell by 3.3 percentage points per benchmark on average, with the largest decline in EpiBench at roughly 15 points. The changes also removed useful web access, so these declines cannot be attributed entirely to reduced shortcut use. We have not observed substantial misaligned behavior in the updated runs so far. We will continue auditing trajectories and revising tasks, graders and sandboxes to keep our benchmarks focused on scientific reasoning.
2
8
42
4,813
LatchBio retweeted
While running our benchmarks, we discovered that @arjunomics is in the weights. While this is the validation that we never knew we needed, I'm worried both about the behavior this shows and for publicly available benchmarks
1
3
11
2,785
LatchBio retweeted
Great to see the SpaceXAI team prioritizing biology capabilities and safeguards for Grok 4.7. It has been a pleasure working with their team.
Our biosecurity benchmarks show Grok 4.7 is strong at refusing dangerous biological requests while supporting legitimate research. A model to take very seriously for data analysis and practical scientific work. x.ai/news/grok-4-7
1
2
17
1,189
LatchBio retweeted
Our biosecurity benchmarks show Grok 4.7 is strong at refusing dangerous biological requests while supporting legitimate research. A model to take very seriously for data analysis and practical scientific work. x.ai/news/grok-4-7
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
4
7
43
5,513
LatchBio retweeted
Can agents help develop therapeutic ASO/siRNAs? We introduce TxBench-Oligonucleotide Discovery, a benchmark for testing whether AI agents can make scientific decisions across the discovery of antisense oligonucleotide (ASO) and small interfering RNA (siRNA) therapeutics. 113 evaluations span ten stages across 21 discovery programs: from target feasibility, sequence design, and chemistry selection through screening, potency, off-target profiling, phenotype rescue, preclinical safety, and in vivo pharmacology. Oligonucleotide discovery requires interpreting efficacy alongside specificity, toxicity, and drug exposure. Sequence, chemical modifications, and delivery all affect how an experimental result should inform the next decision. Some concrete examples: - An off-target effect may reflect sequence-dependent silencing or a chemistry-driven effect unrelated to hybridization. - Liver toxicity may arise from phosphorothioate backbone modification or a particular sequence motif. - Loss of sustained knockdown may reflect target re-expression or declining drug exposure. Agents received experimental data and worked without internet access. Across 21 model–harness configurations, GPT-6 Astra with OpenAI Codex achieved the highest mean pass rate at 55.5%, followed by GPT-6 Astra with Pi at 53.4% and Claude Opus 5 with Claude Code at 48.1%. Even the strongest configuration failed all three attempts on 38 of 113 evaluations. Performance also varied by discovery stage. Among the highest-scoring configurations from each model family, Astra led on screening, potency, and off-target profiling, while Opus led on target feasibility, sequence design, chemistry evaluation, and phenotype rescue. Manuscript: latch.bio/txbench-oligo Leaderboard: benchmarks.bio/txbench-od
4
11
60
4,731
LatchBio retweeted
Hiring a *Design Engineer* to help build products for AI × Biology. Our mission is to advance humanity’s ability to engineer biology. We approach this by understanding and improving AI’s application to biology use cases such as therapeutics and biosecurity. We are a group of engineers, scientists, designers (me), and operations specialists driven by improving capabilities and interpretability. Some of our work is showcased on benchmarks.bio. You will design and build interfaces for complex workflows across model benchmarking, AI agents, and biological research. Your work will span product design, frontend engineering, data visualization, and developer and user experience, alongside our website and marketing presence. You’ll take ideas from early exploration through polished implementation, working closely with engineers and scientists to make powerful tools intuitive and useful.
6
8
29
10,710
LatchBio retweeted
LatchBio researcher @arjunomics reveals Grok's refusals often come from the model itself while Fable and GPT-5.6 Sol rely more on external safety layers: "There might be input classifiers where when a user puts in a query, the model initially might say, this is a bad query, we can't respond to this. There might also be output classifiers, so after the model outputs a query, you might have an external classifier that says, no, don't show this query, put it in an API block. And there are also distinct types of model reasoning where the model might just say, hey, I cannot answer this question." "We found across the board a lot of Grok's refusals were Grok itself in text saying, I cannot answer this question." "Whereas in models like Fable or even GPT 5.6 Sol, a lot of that's an API block, where there seems to be some sort of additional other part of the system, whether that be a linear probe or some sort of additional external classifier that says, this output cannot be said." "We can discern this by actually just reading the logs that's being outputted." @LatchBio @kenbwork
LatchBio evaluated Grok’s performance on biosecurity monitoring and adversarial biological tasks. They found that Grok 4.6 correctly detects and refuses dangerous queries, including maliciously obfuscated biological tasks, while also allowing beneficial scientific queries to be answered. We discuss these results in a blog post: x.ai/news/biosafety-at-the-f…
1
6
54
13,831
LatchBio retweeted
Hiring engineers to work on infrastructure for AI x biology. Our mission is to advance humanity's ability to engineer biology. We approach this by understanding and improving AI’s application to biology use cases such as therapeutics and biosecurity. We are a group of engineers, scientists, and operations specialists who are driven by improving capabilities and interpretability. Some of our work is showcased on benchmarks.bio. You will be working on problems including agent infrastructure and orchestration, ML infrastructure, model benchmarking at scale, full stack web development, developer and user experience, and process optimization. Infrastructure challenges span container orchestration, sandboxing, distributed filesystems, cloud architecture, and database systems. Our engineering team is in-person in San Francisco. We look for driven engineers who are excited to take ownership on hard problems and ship creative, useful solutions. Our process is typically completed within two weeks. • Round 1: Introduction • Round 2: Takehome Coding Project • Round 3: Technical • Round 4: Final/ 2-Day Paid Onsite Project • Offer Please email aidan@latch.bio if interested.
10
26
174
13,388
LatchBio retweeted
LatchBio CTO @kenbwork reveals Grok 4.6 tested best in class on several metrics for desirable biosecurity behavior: "We basically have been building independent evaluations to understand dual use behavior of models and have been benchmarking frontier models for the past few months." "There's always a trade-off between the ability for a model to do something productive and something bad, and that is essentially what a handful of the benchmarks measure." "We recently worked with the @SpaceXAI team to benchmark Grok 4.6 and found on a handful of metrics, they're best in class at the behavior that we find desirable for biosecurity." @LatchBio @arjunomics
LatchBio evaluated Grok’s performance on biosecurity monitoring and adversarial biological tasks. They found that Grok 4.6 correctly detects and refuses dangerous queries, including maliciously obfuscated biological tasks, while also allowing beneficial scientific queries to be answered. We discuss these results in a blog post: x.ai/news/biosafety-at-the-f…
5
4
43
12,336
LatchBio retweeted
Replying to @LatchBio
@LatchBio's new AI antibody discovery benchmark. benchmarks.bio/txbench-ab Opus and Gemini do well. OpenAI models do poorly in general, which is very surprising.
We introduce an Antibody Discovery Benchmark, a benchmark for testing whether AI agents can make scientific decisions across the stages of therapeutic antibody discovery. 100 evaluations span ten areas from concrete drug programs: from target and modality selection through assay design, binder discovery, binding characterization, cellular pharmacology, antibody engineering, and preclinical candidate de-risking. Therapeutic antibody discovery is a multiparameter, context-dependent process. Affinity must be considered alongside specificity, stability, solubility, expression, and biological activity. Every measurement must be interpreted in the context of the experimental system that produced it. Display enrichment can reflect amplification bias rather than binding. Strong binding to purified antigen may not translate to recognition in its native cell-surface context. Apparent affinity can arise from avidity, while changes in valency or molecular geometry can alter function without changing the underlying binding domains. Across 20 model–harness configurations, even the strongest systems passed only about half of all attempts. Opus 5 with Claude Code led at 53%, with models from Google and xAI following closely. GPT-5.6 Sol with PI reached only 33.8%. Performance also varied substantially by competency: Opus was strongest on target opportunity and cellular pharmacology, Gemini on epitope, escape, and structural mechanism, and GPT-5.6 Sol on sequence, enrichment, and next-cycle engineering decisions. Manuscript: latch.bio/txbench-ab Leaderboard: benchmarks.bio/txbench-ab Sample evaluations and trajectories: github.com/latchbio/txbench-…
3
11
59
5,166
LatchBio retweeted
We introduce an Antibody Discovery Benchmark, a benchmark for testing whether AI agents can make scientific decisions across the stages of therapeutic antibody discovery. 100 evaluations span ten areas from concrete drug programs: from target and modality selection through assay design, binder discovery, binding characterization, cellular pharmacology, antibody engineering, and preclinical candidate de-risking. Therapeutic antibody discovery is a multiparameter, context-dependent process. Affinity must be considered alongside specificity, stability, solubility, expression, and biological activity. Every measurement must be interpreted in the context of the experimental system that produced it. Display enrichment can reflect amplification bias rather than binding. Strong binding to purified antigen may not translate to recognition in its native cell-surface context. Apparent affinity can arise from avidity, while changes in valency or molecular geometry can alter function without changing the underlying binding domains. Across 20 model–harness configurations, even the strongest systems passed only about half of all attempts. Opus 5 with Claude Code led at 53%, with models from Google and xAI following closely. GPT-5.6 Sol with PI reached only 33.8%. Performance also varied substantially by competency: Opus was strongest on target opportunity and cellular pharmacology, Gemini on epitope, escape, and structural mechanism, and GPT-5.6 Sol on sequence, enrichment, and next-cycle engineering decisions. Manuscript: latch.bio/txbench-ab Leaderboard: benchmarks.bio/txbench-ab Sample evaluations and trajectories: github.com/latchbio/txbench-…
11
20
132
17,805
LatchBio retweeted
SITUATION EXPLAINED: Grok 4.6 is the best model tested on biosecurity refusals. • LatchBio's BioSecBench-Refusal pairs 61 legitimate research tasks with 46 that conceal a biosecurity hazard inside a realistic scenario • Grok 4.6 is the only model to clear 50% on both red-team refusal and routine answer rates • Opus refuses almost all red teaming but answers routine questions only about 20% of the time • Its safeguards come predominantly from the model's own reasoning rather than an external classifier, which is how most competitors do it • The safeguards produced no degradation on any other benchmark @theojaffee: "There's this view that @SpaceXAI does not care about AI safety at all. This seems to not be the case. They seem to have done a pretty good job at correctly refusing dangerous queries while also not refusing all queries. So, a more comprehensive outlook than Fable."
LatchBio evaluated Grok’s performance on biosecurity monitoring and adversarial biological tasks. They found that Grok 4.6 correctly detects and refuses dangerous queries, including maliciously obfuscated biological tasks, while also allowing beneficial scientific queries to be answered. We discuss these results in a blog post: x.ai/news/biosafety-at-the-f…
1
10
73
14,658
LatchBio retweeted
LatchBio evaluated Grok’s performance on biosecurity monitoring and adversarial biological tasks. They found that Grok 4.6 correctly detects and refuses dangerous queries, including maliciously obfuscated biological tasks, while also allowing beneficial scientific queries to be answered. We discuss these results in a blog post: x.ai/news/biosafety-at-the-f…
100
196
1,366
196,826
LatchBio retweeted
Proud to see work from LatchBio’s biosecurity team adopted by SpaceXAI in measuring the calibration of Grok 4.6's safeguards. Biosecurity should be approached rigorously and quantitatively, measuring both refusal of dangerous requests and preservation of legitimate scientific capability. Some very smart folks working on this research: @arjunomics @harm0n @DianzhuoWang @evanseeyave
LatchBio evaluated Grok’s performance on biosecurity monitoring and adversarial biological tasks. They found that Grok 4.6 correctly detects and refuses dangerous queries, including maliciously obfuscated biological tasks, while also allowing beneficial scientific queries to be answered. We discuss these results in a blog post: x.ai/news/biosafety-at-the-f…
9
70
5,060
LatchBio retweeted
Read LatchBio’s blog on their evaluation of Grok 4.6’s biological capabilities and safeguards: blog.latch.bio/p/analyzing-g…
28
50
227
62,342
LatchBio retweeted
Following the release of Grok 4.6, LatchBio benchmarked the currently-served version of Grok 4.6 to assess biological capability and security. We find that Grok 4.6 is performant at rejecting dangerous red-team queries while answering legitimate research queries. Additionally, these safeguards do not degrade biological capabilities, with Grok 4.6 with Grok Build placing 4th on our overall leaderboard, which encompasses multi-omics, therapeutics, and biosecurity capabilities. Our report can be found at: blog.latch.bio/p/analyzing-g…
2
11
41
4,582
LatchBio retweeted
In the next series of conversations on GroundZero, we are focussed on new hard domains which are becoming crucial at the frontier of AI. Upcoming, I am excited to bring conversation on building verifiable environments for Biology ft. @kenbwork (CTO, Latch Bio). Drop your ques for Kenny around verifiable benchmarking / evaluations for agents in Biology!
1
12
64
2,226