@UCSF facial/reconstructive surgeon. Researching AI for clinical reasoning. Cofounder @memorahealth (acq). Alum @ycombinator w18, @broadinstitute, @harvardmed.

San Francisco, CA
1/ When expert clinicians see similar cases, they often reuse stable mental templates, known as diagnostic schemas. I built a way to test whether LLMs do the same. (Below: my study group thinking through a cardiogenic shock case during med school)
3
5
22
6,160
The font/color choice here by Spotify is a little too on the nose
1
1
8
467
Catch me listening to this on my morning commute to log into CPRS.
2
13
1,946
✌🏽ICML
1
18
2,006
Excited to try this (especially since OE has full text data)! For the last couple months I’d been routing at least half of my clinical questions into the CLI with my evidence-grader skill for the exact same reason; it’s hard to gauge the quality of evidence in these answers short of reading the papers myself or uploading their full text into a model and asking for the appropriate study appraisal tool.
Some clinical questions have thirty years of randomized trials behind them. Some have four case reports. Answers on OpenEvidence now carry a grade for how much certainty the underlying evidence can support, using the same GRADE criteria Cochrane and the WHO apply to guidelines. Study design, agreement across trials, precision, and whether the studied population resembles the patient in front of you. @FierceHealth got the first look at EvidenceGrade. Link in the reply below.
4
741
Heading to Seoul for ICML! Presenting a Spotlight paper on Saturday at the Structured Data for Health workshop. If you work on AI for health/bio, AI safety, uncertainty calibration, or CoT faithfulness come say hello or DM me. 📍July 11, COEX Hall D2, 9:50-10:50AM
1/ When expert clinicians see similar cases, they often reuse stable mental templates, known as diagnostic schemas. I built a way to test whether LLMs do the same. (Below: my study group thinking through a cardiogenic shock case during med school)
11
1,170
1/ When expert clinicians see similar cases, they often reuse stable mental templates, known as diagnostic schemas. I built a way to test whether LLMs do the same. (Below: my study group thinking through a cardiogenic shock case during med school)
3
5
22
6,160
9/ If the consistency we'd expect isn't there across similar cases, a deeper question is whether anything schema-like is represented *inside* the model at all, or whether each diagnosis is rebuilt from scratch. That's the mech interp problem I'm studying next.
1
493
10/ I'll be at ICML presenting this at the Structured Data for Health (SD4H) workshop on July 11. If you work on clinical AI evals, CoT faithfulness, or graph representations and you'll be there, find me or send a DM. I'd love to chat and am actively looking for collaborators! Paper: arxiv.org/abs/2606.29876 Session: July 11, COEX Hall D2, 9:50-10:50AM (structureddata4health.github…)
2
5
502
Nisarg Patel, MD retweeted
Come work with me! I am hiring a postdoc to work with me on a set of projects studying the deployment of clinical AI. (Link to apply below)
1
2
6
2,721
In medicine, we’re already seeing startups expand into clinical decision support tools with read-access to the entire health record, which are often *decades* long per patient. Compressing not only text, but imaging, labs/trajectory, and reasoning in the least lossy way possible is critical, and representation-level tools may be the way to do it! Check out Vishnu’s work on the subject.
I've been researching how to compress coding agent context at the representation level, working directly with models' internal vectors instead of summarizing text. So far, I've hit 2x+ compression while preserving most factual recall.
2
34
6,848
Seeing a lot of bitter lesson/scale > specialized model tweets about this study, but short of OE/UTD intentionally buying weaker models, this result is more of a *harness* issue than a model one. MedQA score ranges across models were within 9 points (so less likely a base knowledge gap); the gap between gen vs clinical models widened on the real world clinical question answers, where the problems were mostly in omission and disorganized outputs rather than wrong facts.
For medical information, general AI frontier models (Google, OpenAI, Anthropic) outperformed specialized @openevidence and @UpToDate as assessed by 12 US clinicians, randomized and blinded to which model and extensive testing/benchmarks. This was not anticipated. @NatureMedicine nature.com/articles/s41591-0…
7
2
23
6,352
Me teaching every surgeon I meet how to use ChatGPT for Clinicians / OpenEvidence.
Forward Deployed Clinician (FDC)
7
4
73
13,397
Immensely proud.
I wanted my first public video to be learnings from the past 7.5 years here at @AppliedInt. Hopefully it’s useful to the next set of builders 🐻
3
1,640