Today, we are incredibly excited to present LiteMol-1, our very first foundation model from LiteFold.
Structure-based models like BoltzGen, O-Design, and RFdiffusion have become the de facto standard for designing biomolecules. However, they are expensive to run at scale. In many campaigns, we have to generate tens of thousands of designs and rigorously filter them down to a handful of top candidates.
More importantly, most molecule design systems are primarily optimized around binding. But what about everything else that makes a molecule a useful therapeutic: ADME, toxicity, drug-likeness, membrane permeability, synthesizability, selectivity, and more?
That’s where we introduce LiteMol-1, our first Multi-Molecule Foundation Diffusion Language Model, pre-trained from scratch. With a single set of weights, LiteMol-1 can conditionally generate small molecules, peptides, cyclic peptides, depsipeptides, peptides with ncAAs, macrocycles, and PROTACs.
You can generate molecules unconditionally or condition generation on a protein target sequence. You can in-paint molecules, preserve parts of an existing scaffold, generate non-canonical peptides, cyclize designs, and continuously edit and regenerate them.
We also introduce a Monte Carlo Tree Search framework for multi-objective molecular design. Instead of trying to make one molecule perfect at everything, the search keeps a set of promising molecules, each with different trade-offs across properties. This matters because molecular design is fundamentally not a single-objective optimization problem.
Finally, and probably the part we are most excited about: this is a model for Agents. LLMs are not particularly efficient interfaces for repeatedly reasoning over thousands of large PDB/CIF files inside an AutoResearch loop. Sequence space is different. SMILES and molecular sequences are compact, editable, and much easier for an agent to inspect, compare, modify, and reason over.
So we treat LiteMol-1 as an infinite molecular canvas.
The model generates possibilities. The agent takes inspiration from them, evaluates them, edits them, optimizes them, and generates again. The AutoResearch loop continues until it finds candidates that satisfy the given design objectives; while bringing in expensive structure prediction, docking, or simulation only when they are actually needed.
Our evaluations across peptide and small-molecule generation show that LiteMol-1 is competitive with, and in several settings on-par with or better than, frontier structure-based and sequence-based models, while operating at a fraction of the generation cost and time.
And this is only the first step. Check out our technical research blog post in the comments to learn more about LiteMol-1.