Hugging Science retweeted
The SEPIQ challenge has released a large dataset of +1.8 million VHH sequences to make progress on one of computational biology's toughest open problems. Join the challenge and help build the next generation of computational tools for predicting antibody–antigen interactions: huggingface.co/spaces/sepiq-…
only a few people seem to know about the incredibly rich data on antibodies for cancer that was released as part of the SEPIQ challenge absolute gold data that may have the answer to breast cancer in it
2
24
136
12,832
Hugging Science retweeted
If you want to predict next weather - wind, temperature, etc... - you can now do it yourself! You can run and compare AI based weather forecasting models (Atlas from @nvidia, AIFS from @ECMWF, Aurora from @Microsoft) on our demo space: huggingface.co/spaces/huggin… Thank you so much for making the weights available on @huggingface !
9
33
327
20,334
Hugging Science retweeted
support for mmCIF, PDB, fasta, and fastq just hit the `datasets` library this should massively accelerate 🧬🧫🧑‍🔬 bio workflows on @huggingface huge kudos @b_azarkhalili & come coming soon 📈
3
4
27
2,441
Hugging Science retweeted
Can you predict next week's weather yourself? We just wrote a blog post that explains how! ⛅️🌍🌧️ Running weather forecasting required having a supercomputer and CS/physics Phd. Now, AI-based models need fewer resources and are available to anyone on @huggingface. So, prediction the weather should be accessible to all. But only a few labs can actually run these models because there is no standardized documentation guiding you from the model weights to an actual forecast. Together with @EarthmoverHQ, we wrote a blog post and a demo walking you through that: huggingface.co/blog/hugging-… The blog covers the code for fetching initial conditions from the Earthmover Marketplace, building a model's input batch, running inference and comparing the forecast against historical climate data. You can also run Aurora (Microsoft), WeatherNext2 (Google) and AIFS (ECMWF) on our demo space. This is a first step to make AI-based weather forecasting models more accessible, and we'd love to build more of that. Tell us what's missing or your feedback in the blog comments!
8
25
124
19,739
Hugging Science retweeted
🌱🌱🌱 a HUGE step towards designing to food of tomorrow -- from the DNA up 🧬🧬🧬 🙏 BOTANIC-1 just dropped, a family of models to do just that (beating all prev models) All open on @huggingface 🤗 + more useful than Millennium Problems More in 🧵
7
22
108
6,864
Hugging Science retweeted
Millennium Problems are just the beginning—to realize epic science we need to be pushing across many domains (most of which have no verifiable rewards). Come work with us on these hard problems (together with @AnthropicAI @futurebio_xyz and @HackIterate) this October.
6
12
108
6,056
Hugging Science retweeted
There is a child with a rare disease who is currently suffering and struggling to manage his symptoms. Rare as this is, you can directly help him. Today we are launching the "Rare Disease, Real Kid" Hackathon, and there are $50,000 in prizes from @AnthropicAI and @awscloud. We (@huggingface & @Sagebio) are helping this child open his genome and clinical data to the community, so that we can find what's caused his disease and what currently-approved drugs could help him. I doubt I need to motivate this much further or explain how rare it is for a family to share their child's genome and clinical data, but if you're not sure, consider this: Until very recently, it wasn't feasible for patients like this to get treatment because their disease was so rare that the economics could never justify the investment. Now, as we've seen, people with rare diseases are starting to be able to find the answers themselves (with the help of AI tools, cheaper sequencing, etc). This kid is not able to do that for himself and neither are his parents, so we're asking you for help. Both for this kid and to prove that it's possible for everyone else suffering from a rare disease. More details in 🧵. sagebio-rare-disease-real-ki…
124
495
3,189
664,709
Hugging Science retweeted
Announcing Sim2Science: ML with Imperfect Scientific Models @NeurIPSConf! Submission deadline: 29 August 2026, 23:59 AoE Workshop: 12 or 13 December 2026, Paris, France 5-page workshop papers & 2-page Tiny Papers via OpenReview sim2science.com
3
11
1,518
Hugging Science retweeted
Seems like no one's noticed the 80TB of astrophysics data from 30+ sources that just dropped on @huggingface. ...and you only need ~4GB of RAM to load it. We're talking over 80TB of galaxy imagery taken across the spectrum, spectra of galaxies and stars, time series of variable stars, and a whole zoo of assorted measurements and physical data. And all of it can now be wrangled on your laptop, thanks to Multimodal Universe's just released cross-matching. SDSS x Gaia means you can match 800k objects against 122M objects, and it never climbs above ~4GB of RAM. Huge congrats to @smith42mike for leading this and making the world of astro accessible to probably 10,000x more people. Let's discover some shit
40
222
1,418
122,066
Hugging Science retweeted
Together with UC Berkeley we are announcing the laser phase plate - a breakthrough in atomic resolution imaging. This is the brightest continuous wave laser in the world, 100 million times the intensity of the surface of the sun. Phase contrast plays an important role in microscopy, but it was thought close to impossible for electron microscopy, where it would require interfering with an electron beam. Holger Mueller and Robert Glaeser proposed exactly this using a standing wave laser. It has taken over 15 years to make this a reality. Biohub partnered with UC Berkeley and Mueller to support this work and to engineer and build the technology. Contrast has been the critical barrier to achieving atomic resolution imaging of the cell. In cryo-electron tomography, a cellular imaging technology that uses electron microscopy, the low contrast makes it impossible to resolve anything but the largest proteins within their cellular context. The laser phase plate removes that barrier. With advances in AI this breakthrough in contrast will start to open up a new frontier in structural biology, that will allow us to see the molecular machines of the cell, and how they assemble into far more complex and dynamic systems, and understand how they work.
93
565
3,952
661,905
Hugging Science retweeted
Use nano-scGPT to compute SOTA cell embeddings in just 6 lines of code, now available on @huggingface @huggingscience. nano-scGPT isn't an inference wrapper, it's scGPT reimplemented from scratch. Finetuning and pretraining are coming next! Let me know what models/tasks you'd like to see next.
Just 6 lines of code to run scGPT. I built nano-scGPT, the simplest implementation of single-cell foundation model scGPT (by @BoWang87, @HAOTIANCUI1), inspired by @karpathy's nanoGPT and @ChrisHayduk's minAlphaFold2. Written in pure PyTorch (~270 loc for model, ~230 loc for tokenizer) with minimal dependencies. nano-scGPT produces embeddings numerically equivalent to the original while running 1.38x faster, mostly from a clean forward pass plus torch.compile. The hope is to get more people up to speed with bio/cell AI/ML, which is arguably the most exciting AI/ML field right now, and to make bio foundation models like scGPT much more accessible to run, understand, and tinker with!
1
1
3
199
on hugging science: kimina-prover ✍️ rl-trained model that writes formal lean 4 proofs for olympiad-level maths. machine-checkable, not just "showing working". 7b distilled preview, open on the hub. huggingface.co/AI-MO/Kimina-…
1
151
Hugging Science retweeted
Also big thanks goes to @huggingface for enabling open source AI and AI4science. we have been using HF for the longest time and our family of highly performant fMRI foundation models (CortexMAE) is available on HF along with the training dataset. @cgeorgiaw
1
2
22
10,898
Hugging Science retweeted
We’re excited to share the full binder design protocol. Check it out here: github.com/Biohub/esm/blob/m…. The notebook includes support for @modal to easily scale up binder generation. Give it a try and let us know how it works! You can read more about ESMFold2, ESMC, ESM Atlas, and the full results in the paper here: biohub.ai/papers/esm_protein….
3
25
82
23,013
Hugging Science retweeted
huge for getting medical data ACTUALLY used by the machine learning community All publicly funded science should be open source 🔥🔥🔥
Hugging Face is the home for AI & ML across every domain, including biomedical! The @NIH just added the @huggingface Hub to its official list of Generalist Repositories for data sharing. NIH-funded? You can point to the Hub in your data sharing plan 🤗
6
13
93
18,276
Hugging Science retweeted
Hugging Face is the home for AI & ML across every domain, including biomedical! The @NIH just added the @huggingface Hub to its official list of Generalist Repositories for data sharing. NIH-funded? You can point to the Hub in your data sharing plan 🤗
5
23
75
28,172
Hugging Science retweeted
Opus 4.8 just dropped and I ran it through our CAD tasks. 4.6 → 4.7 → 4.8 side by side. The results are unexpected!
194
188
3,478
709,899
on hugging science: mattergen ⚛️ generative ai for materials. you give it a target property, it proposes novel inorganic crystal structures to match. inverse design instead of screen-and-pray. built for energy, catalysis and functional materials research. weights on the hub.
1
7
449
Hugging Science retweeted
today was a massive day for protein engineering. esmfold2 dropped—next gen of the esm series, fully open on @huggingscience. 1.1 billion predicted structures, 6.8 billion sequences. 800m more entries than the alphafold db, and reportedly edging out alphafold3 on protein complexes, including antibody–antigen binding. alongside it: the new esm atlas. a huge expansion of known protein space, heavy on metagenomic sequences from soil, ocean, and the parts of biology that have been least characterised (until now!!) and if that weren't enough, litefold dropped the fineweb of proteins, so every major protein database (pdb included) aggregated, cleaned, and made plug-and-play in one place. these are the releases that push the whole field forward, and the pace of open science right now is almost motion-sickness inducing all of it on huggingscience.co (and ofc @huggingface)
9
72
342
36,646