Bioinformatics Senior Scientist @rxbiologics. On a mission to improve drug discovery using machine learning (all views expressed are my own)

Replying to @CodeWisdom
"Programming isn't about what you know; it's about what you can copy and paste from stack overflow" - someone in a hurry
4
5
33
William Steele retweeted
🔓🧬First big unlock from vibe science-ing: rapid access to publicly available datasets. Sounds basic. It isn't. If you've ever tried to pull raw data from a paper you care about, you know: the metadata is a mess, the supplementary tables are unstructured, the file formats don't match, and by the time you've got it working you've lost half a day. For people without strong bioinformatics skills, it's often a dead end entirely. 1/🧵
7
19
117
12,921
William Steele retweeted
I'm rebuilding AlphaFold2 from scratch in pure PyTorch. No frameworks on top of PyTorch. No copy-paste from DeepMind's repo. Just nn.Linear, einsum, and the 60-page supplementary paper. The project is called minAlphaFold2, inspired by Karpathy's minGPT. The idea is simple: AlphaFold2 is one of the most important neural networks ever built, and there should be a version of it that a single person can sit down and read end-to-end in an afternoon. Where it stands today: - ~3,500 lines across 9 modules - Full forward pass works: input embedding → Evoformer → Structure Module → all-atom 3D coordinates - Every loss function from the paper (FAPE, torsion angles, pLDDT, distogram, structural violations) - Recycling, templates, extra MSA stack, ensemble averaging — all implemented - 50 tests passing - Every module maps 1-to-1 to a numbered algorithm in the AF2 supplement The Structure Module was the most satisfying part to build. Invariant Point Attention is genuinely beautiful — it does attention in 3D space using local reference frames so the whole thing is SE(3)-equivariant, and the math fits in about 150 lines of PyTorch. What's next: - Build the data pipeline (PDB structures + MSA features) - Write the training loop - Train on a small set of proteins and see what happens The repo is public. If you've ever wanted to understand how AlphaFold2 actually works at the level of individual tensor operations, this is meant for you. Repo: github.com/ChrisHayduk/minAl…
58
250
2,258
83,387
William Steele retweeted
In January, @jonhoo, @jjgort, and I returned to @MIT_CSAIL to teach Missing Semester, a class on topics missing from most CS programs—tools and techniques that everyone should know, like Bash, Git, CI/CD, and AI tools. Today, we’re releasing the course for free online!
16
209
1,472
90,376
William Steele retweeted
🎁 We have a gift for you! You've heard about skrub and would like to discover more? Or you never heard about it, but struggle with data preprocessing? 📽️ Riccardo Cappuzzo did an awesome video at PyData that has been recorded: you can have a look here 👉 eu1.hubs.ly/H0qDx3G0
3
7
1,464
William Steele retweeted
Hey, all the AI x Bio companies. I have an idea. if your agent can read this paper nature.com/articles/s41586-0… download the data, and replicate the figures, I am going to subscribe.
10
29
311
32,056
William Steele retweeted
2
5
179
41,542
William Steele retweeted
Exciting preview of OpenFold3 has just been released! Congrats on the first milestone to the OpenFold team! And I must say I'm eagerly awaiting the full release with ablations and scientific insights. This is something that other releases in our field just do not do, apart from OpenFold. But I think this is what pushes the community forward scientifically/in terms of understanding!
🔥 It's here: OpenFold3 is now live. THE open-source foundation model for predicting 3D structures of proteins, nucleic acids & small molecules. This is where the future of drug discovery and biomolecular AI lives. Built by @open_fold. Hosted on @huggingface. 👇 more
7
55
6,085
William Steele retweeted
If you build more housing, it becomes cheaper. Might be shocking to some people, but supply and demand never fails.
88
92
1,154
226,852
William Steele retweeted
Our new deep protein structure representation layer, BioBlobs is out on arXiv. Protein function operates by coordinating cohesive substructure modules shaped by evolution. Current protein representation methods break proteins down into rigidly shaped and sized blocks. Would you model a bicycle by breaking it into uniform blocks? We built a differentiable graph partitioning model which learns to identify these cohesive 3D modules (blobs) and use them to build protein embeddings. Our embeddings showed significant improvement when placed on top of GVP-GNNs on 3 protein function prediction tasks. Blobs are assigned importance scores which can help us better understand the mechanisms of protein function. Thank you to Allen (Xin) Wang for his outstanding effort in leading this project and making the first paper of my lab a reality. Preprint: arxiv.org/abs/2510.01632 Code: github.com/OliverLaboratory/…
4
20
146
7,063
William Steele retweeted
Imagine how silly it is that pipetting "robots" still has no clue what they are doing. No live deck checking,no liquid monitoring etc.
We built a robot brain that nothing can stop. Shattered limbs? Jammed motors? If the bot can move, the Brain will move it— even if it’s an entirely new robot body. Meet the omni-bodied Skild Brain:
1
1
1
317
Why is it that to build the simple version of something, you always seem to have to build the complicated one first🤔?
24
I never get why K-means is more popular than Hierarchical clustering. When do you ever know the ideal number of clusters up front?
29
Just uploaded a new post on NumPy indexing tricks. NumPy is an awesome python library for numeric computation, but it can be confusing. Hopefully this post can save people a few headaches! #Python wjs20.github.io/2025/07/10/n…
26
Fantastic book on using the shell effectively. I've been using the shell for a while and I'm learning new stuff! effective-shell.com/
33
If ChatGPT says 'lets unpack this carefully' one more time I might actually cry.
1
22
Like I asked you to remind me how make a line plot in matplotlib, what is there to unpack...
1
19
William Steele retweeted
Chai Discovery (@joshim5 @chaidiscovery) just released Chai-2, a model for "zero-shot antibody design in a 24-well plate." The most existing result? Chai-2 successfully generated binders for 50% of the 52 targets they assayed (100x better than SOTA)! Some thoughts below:
2
47
270
44,348