CEO @SophontAI | Founder @MedARC_AI | PhD at 19 (2023) | ex Research Director Stability AI | Biomed. engineer @ 14 | TEDx talk➡bit.ly/3tpAuan

NEW RESULTS: Our researcher @benjamin_warner has been diligently working on updating the Medmarks leaderboard (evaluating model's medical capabilities) with some of the latest mid-size open-source models. Some interesting results, read to find out more:
This year, we released Medmarks v1.0, our open benchmark suite + leaderboard of LLM medical capabilities. Today, we release additional results for some of the latest mid-size open-source LLMs including Gemma, Qwen, Muse, Nemotron. We find Gemma 4 31B leads in this size class. Link: sophont.med/blog/medmarks-su…
1
5
23
4,474
Tanishq Mathew Abraham, Ph.D. retweeted
Medmarks accepted at NeurIPS!
I'm excited to share Medmarks was accepted to the NeurIPS datasets and benchmarks track! 🥳 See you at Sydney!
1
3
701
I'm excited to share Medmarks was accepted to the NeurIPS datasets and benchmarks track! 🥳 See you at Sydney!
We're excited to release Medmarks v1.0 + a technical report! This is an update to our Medmarks benchmark suite, the largest open-source automated suite for evaluating the medical capabilities of LLMs. We added 10 benchmarks (20→30) and 15 models (46→61) to the leaderboard!
8
8
66
4,799
The cybersecurity risks of the AI age are enormous!! @Shalev_lif and @RomiLifshitz have been thinking about this problem deeply and have written a stellar report about these unique challenges and what could be done to secure the path to superintelligence.
We are approaching cyber-superintelligence. But our systems aren't ready. We must secure our model weights and infrastructure against three new threats: sabotage, escape, and theft. There is a way forward. Secure Acceleration: A Cyberdefense Strategy for Superintelligence
3
13
2,867
Tanishq Mathew Abraham, Ph.D. retweeted
We are approaching cyber-superintelligence. But our systems aren't ready. We must secure our model weights and infrastructure against three new threats: sabotage, escape, and theft. There is a way forward. Secure Acceleration: A Cyberdefense Strategy for Superintelligence
80
85
570
135,501
bruh they turned glasses into an FDA-cleared hearing aid...
1
1
70
4,952
someone is holding up their glasses to record at Meta Connect rofl
4
1
66
5,856
Okay my take on the Anthropic bio announcement: It's very cool and people should be excited! but it's also very very very preliminary! They discovered a new enzyme system but the basically don't know anything about how it works. The only real lab experiment they've done is just express it in E coli... Imo this announcement speaks to a broader idea of how LLMs can be used to accelerate biological discoveries: automating the process of finding something unexpected and unusual in biological data to scale up discoveries and hypothesis generation.
6
5
74
7,506
It's kinda wild to see LLMs having this sort of "a-ha"/lightbulb moment in the biological setting!!
ANTHROPIC ANNOUNCES BIOLOGY LAB AND NEW DISCOVERY In Spring of 2026, Anthropic started a research group to accelerate biological discoveries. Today they announced their first result: Claude Mythos 5 autonomously discovered a novel enzyme system that is associated with an array of DNA repeats, a pattern reminiscent of CRISPR. The key insight here behind what Anthropic is trying to do: A lot of biological discoveries come from scientists noticing something unusual and unexpected. LLMs could potentially automate this manual process, scaling up the amount of potential discoveries. Here, the researchers focused Claude's attention onto reverse transcriptases (RTs). These are enzymes that copy RNA into DNA. RTs are used for many biotechnologies from DNA synthesis to genome engineering. So discovering novel RTs can be very valuable! Claude was made to search autonomously through a massive database of DNA sequences for interesting new examples of RTs. After 21 hours spent by roughly 950 agents using 210 million tokens, one agent spotted something unusual: a repeating pattern of DNA sequences that occurs next to the gene for an odd-looking RT. After further analysis it was convinced that it had found a new biological system! The system it found, ART, is found mainly in bacteriophages (viruses that attack bacteria). The researchers did some preliminary lab experiments where they expressed the ART system in E. coli bacteria and did further RNA sequencing. However, there are still many many open questions of how exactly this ART system works and what its function is. Link: anthropic.com/news/claude-di…
2
7
86
8,192
ANTHROPIC ANNOUNCES BIOLOGY LAB AND NEW DISCOVERY In Spring of 2026, Anthropic started a research group to accelerate biological discoveries. Today they announced their first result: Claude Mythos 5 autonomously discovered a novel enzyme system that is associated with an array of DNA repeats, a pattern reminiscent of CRISPR. The key insight here behind what Anthropic is trying to do: A lot of biological discoveries come from scientists noticing something unusual and unexpected. LLMs could potentially automate this manual process, scaling up the amount of potential discoveries. Here, the researchers focused Claude's attention onto reverse transcriptases (RTs). These are enzymes that copy RNA into DNA. RTs are used for many biotechnologies from DNA synthesis to genome engineering. So discovering novel RTs can be very valuable! Claude was made to search autonomously through a massive database of DNA sequences for interesting new examples of RTs. After 21 hours spent by roughly 950 agents using 210 million tokens, one agent spotted something unusual: a repeating pattern of DNA sequences that occurs next to the gene for an odd-looking RT. After further analysis it was convinced that it had found a new biological system! The system it found, ART, is found mainly in bacteriophages (viruses that attack bacteria). The researchers did some preliminary lab experiments where they expressed the ART system in E. coli bacteria and did further RNA sequencing. However, there are still many many open questions of how exactly this ART system works and what its function is. Link: anthropic.com/news/claude-di…
5
15
70
12,311
Great to see Anthropic reporting biomedical image analysis capabilities...
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
1
2
22
3,554
rofl got downgraded asking the Opus 5.5 to cure cancer 🤣
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
11
4
115
13,268
will it be a banked reset? please be a banked reset 😭
Ladies and gentlemen... start... your... ENGINES. We are almost Tuesday and I promised a reset for Tuesday. Among some other things. See you soon.
16
2
145
22,737
Tanishq Mathew Abraham, Ph.D. retweeted
We are in the process of evaluating more models for the next release of Medmarks, including Astra and Fable, but first, a teaser.
This year, we released Medmarks v1.0, our open benchmark suite + leaderboard of LLM medical capabilities. Today, we release additional results for some of the latest mid-size open-source LLMs including Gemma, Qwen, Muse, Nemotron. We find Gemma 4 31B leads in this size class. Link: sophont.med/blog/medmarks-su…
2
2
9
1,271
Tanishq Mathew Abraham, Ph.D. retweeted
This year, we released Medmarks v1.0, our open benchmark suite + leaderboard of LLM medical capabilities. Today, we release additional results for some of the latest mid-size open-source LLMs including Gemma, Qwen, Muse, Nemotron. We find Gemma 4 31B leads in this size class. Link: sophont.med/blog/medmarks-su…
3
3
17
7,335
honestly me finishing my phd (and with it, my academic education) in spring 2023 was me catching the last chopper out of Nam
14
137
14,422
Just discovered CancerBench was also popular on Threads lol
CancerBench: the frontier model cancer cure benchmark. AI lab CEOs keep talking about curing cancer, so I made a benchmark. One metric: how many types of cancer has your model cured? All models are currently tied at zero. It’s time to hillclimb! cancerbench.com
2
1
45
6,899
I have refrained from commenting on this guy in the past. But this time I have to be clear: this guy has absolutely zero clue what he's talking about. Here, he is completely and utterly wrong about AI. All he is good at doing is sounding confident and incorrectly using technical jargon, and sounding like he's enlightening you on the secrets of the world. I hope no one in my audience is listening to him.
Professor Jiang Explains Why AI Isn’t Real “I guarantee you it's being manipulated by humans somewhere in India.” (Via Jack Neel)
Community note
Professor Jiang's claim that AI responses are manipulated by humans in India is false. LLMs like ChatGPT generate text autonomously via next-token prediction on neural nets after training; humans aid data/RLHF but not live outputs. techradar.com/computing/arti… arstechnica.com/science/2023/0…
64
27
653
100,522
Many mathematicians are complaining that AI proofs are incomprehensible and therefore does not further our understanding of mathematics. To me this seems like a relatively tractable problem for the AI labs to solve, no? Like couldn't you have some sort of reward model/judge that measures how easy to understand a proof is to understand and use that to post-train or guide the models? Perhaps I am oversimplifying this... if so, please explain why this would be hard?
123
26
387
48,795
This is why the channel is called 3blue1brown...
I probably watched every single @3blue1brown video from Grant Sanderson (big fan) and I am noticing this for the first time 🤯
19
6
506
45,087