QCell is on the cover of AI for Science! 🧬
Building general-purpose ML force fields for biomolecules requires diverse, high-quality QM data covering all four major biomolecular classes. Unlike materials, complexity here arises mainly from conformational space within a limited set of recurring building blocks. Though carbohydrates, lipids, and nucleic acids account for around 40% of cellular dry mass, they remain underrepresented in QM training data.
QCell is built to close that gap, contributing 525k new QM calculations (2–402 atoms) for these three classes, plus relevant solvated ion clusters and dimers, extending coverage beyond small molecules and proteins. Combined with existing datasets (QM7-X, GEMS, and QCML), it brings the collection to 41M molecular systems, all computed at a consistent, non-empirical hybrid PBE0+MBD(-NL) level. We validate structural diversity and train an MLFF that reaches an FMAE of 1 kcal/mol/Å across most subsets.
We hope QCell will be useful for developing more general and accurate ML force fields (ours will be out soon!), and will inspire efforts to close the remaining (still wide) gaps across the biomolecules expressed in living cells.
Thanks to
@sciencebrush for the design!