what can a deeply phenotyped human cohort tell us, and what can today’s AI models do with that information?
we built PhenoBench around the Human Phenotype Project: 90 tasks across 15 clinical domains, using clinical, imaging, molecular, and wearable data.
having those tasks let us ask a few things: which measurements actually help? and do more capable models make better use of them?
on these tasks, tabular foundation models performed better overall, but the gains over ridge regression were small. the LLM results felt like a useful example of the “jagged frontier”: they did relatively well on some tasks and poorly on others. models trained on the same input fields generally did better.
btw, most of the work was figuring out what to ask and how to tell whether a model’s answer was any good. that took much longer than running the models. having done that, trying another model or measurement is a lot easier.