Assistant prof @Berkeley_EECS, core faculty @UCJointCPH. Developing ML methods to study health & inequality. bsky.app/profile/emmapierson…

New paper in Nature Communications! We show you can map out urban flooding by using vision-language models to detect floods in large-scale street scene datasets. In New York City, our method identifies flooded neighborhoods, home to 100k people, that current methods miss. 1/
2
14
54
7,549
Thanks for the thoughtful engagement, @EricTopol! I agree AI has powerful applications to cancer and discuss several in the piece. And if that were the only impact of stronger generalist AI, I'd root for it to arrive ASAP! 1/2
2
1
11
877
New piece in @TheAtlantic! We always hear that AI will cure cancer, and I would immediately benefit if it did. But I argue that racing ahead on generalist AI models creates unclear benefits for cancer that are outweighed by broader societal harms. (Gift link in next post)
7
30
153
52,017
Our lab, within the Berkeley EECS department, is hiring a postdoc! More info and quick application form: forms.gle/41tTVesNqtz33R838 Apply by May 1! Please reshare :)
2
29
118
34,963
Now out in Nature Communications - we have released a migration dataset that is - 4000x more granular than existing public data - highly correlated with Census data - being used by >100 academic, govt, and non-profit teams all over the world See @gsagostini's thread!
Our paper “Inferring fine-grained migration patterns across the United States” is now out in @NatureComms! We released a new, highly granular migration dataset. 1/9
2
2
13
7,116
New paper in @JAMAInternalMed: Patient race is widely used in medical algorithms...but it's unclear how patients feel about this. We conduct the first nationally-representative survey to find out, producing 4 key findings with practical implications - see James' thread!
1/ We’ve heard so much from researchers and academics on the tradeoffs of using race in medicine. But what do patients want? We conducted the first nationally representative survey to systematically study these preferences, now out in @JAMAInternalMed:
1
4
16
7,310
ICE's actions are immoral, illegal, and extremely unpopular. And their budget has skyrocketed. They should not keep receiving such a large share of our tax dollars.
Reminder: DHS legislation that funds ICE at current levels without policy restrictions is in the Senate and needs 60 votes to pass. Deadline is next Friday.
4
34
4,469
See the paper for many robustness checks and discussion of nuances! Our finding persists when using alternate outcomes, statistical models, subsets of the data, and controls satisfying the criteria above.
1
395
New @ScienceAdvances paper proposing a simple test for bias: Is the same person treated differently when their race is perceived differently? Specifically, we ask: is the same driver likelier to be searched by police when they are perceived as Hispanic rather than white? 1/
1
8
27
11,003
Our lab @Berkeley_EECS is recruiting PhD students! We develop ML methods for the health + social sciences in order to build a fairer, healthier world. Apply to @Berkeley_EECS or @UCJointCPH and mention my name in your application! More info: people.eecs.berkeley.edu/~em…
10
106
397
27,410
San Francisco after a storm.
21
1,085
🚨 New postdoc position in our lab @Berkeley_EECS! 🚨 (please retweet + share with relevant candidates) We seek applicants with experience in language modeling who are excited about high-impact applications in the health and social sciences! More info in thread 1/3
7
52
165
38,577
The #NoKings protest we attended was peaceful and passionate. Cars honking constantly in support, American flags and California flags and pride flags, young children and old folks, a woman in king's robes, a little boy with a "bad DOGE much wow" sign.
3
3
44
3,839
Migration data is critical in health/environmental/social sciences. We're releasing a new dataset, MIGRATE: annual flows between 47 billion pairs of US Census areas. MIGRATE is: - 4600x more granular than existing public data - highly correlated with ground-truth data 1/2
1
4
27
6,416
We have a new method, HypotheSAEs, for identifying *interpretable text features that predict a target variable* (aka hypothesis generation). What features of a headline predict engagement? What features of a clinical note predict whether a patient will develop cancer? 1/
1
6
73
7,338
New in @Nature: @leah_pierson and I argue that philanthropic funders should shield science from cuts the Trump administration may make to climate science, infectious disease, etc. Free link: rdcu.be/d6aul Longer version on my website: shorturl.at/muwts
2
13
1,662
Our article on using LLMs to improve health equity is out in @NEJM_AI! 85% of equity-related LLM papers focus on *harms*. But equally vital are the equity-related *opportunities* LLMs create: detecting bias, extracting structured data, and improving access to health info.
5
51
217
47,884
Twitter this morning is showing me posts about an MMA fighter and some dude who claims he has "the best biomarkers in the world". Let's see whether other platforms can be better. Info in picture. Will likely still post here occasionally.
1
19
2,221
Sun setting behind the UN building.
16
1,643
Please retweet: I am recruiting PhD students at Berkeley! Please apply to @Berkeley_EECS or @UCJointCPH if you are interested in ML applied to health, inequality, or social science, and mention my name in your app. More details on work/how to apply: cs.cornell.edu/~emmapierson/
7
191
561
69,160
Onward to @iclr_conf, fueled by rainbow bagels. Good luck to all on submissions!
1
1
36
5,124
A bunch of other key details in the paper. Our finding of bias is robust across different outcomes, statistical models, and controls satisfying the four criteria mentioned above. 5/
1
2
1,009
New paper on race adjustments in clinical algorithms in PNAS! Joint work with the wonderful @annalzink and @oziadias - see Anna's detailed thread below. Anna will be on the job market this year - check out her other great work at zinka88.github.io/!
There are good arguments for removing race from medical algorithms, but there may be unintended consequences. Our PNAS paper finds that race-blind algorithms can *worsen* racial inequalities, bc they can't adjust for racial disparities in data quality. shorturl.at/7ugW5
7
55
9,078
I love watching thunderstorms roll in on summer afternoons (this time lapse is taken over about 2 hours).
1
3
58
6,370
It's been a wonderful 3 years at @cornell_tech. I've loved our island campus from the moment I set foot on it, and will always be grateful to the incredible mentors + collaborators who made me feel lucky every day to have this job.
2
46
2,919
My partner, @serinachang5, will also be joining the Berkeley faculty! We've taken on so many challenges together, from modeling COVID 🦠 📈to dating long distance ✈️🏳️‍🌈; I’m looking forward to taking on this new challenge as well. 2/
7
2
134
5,243
Life update: I'll be joining the @Berkeley_EECS faculty in Jan 2025! I'll also be part of the Computational Precision Health program. I'm excited to continue our work using ML to improve health + social equity at Berkeley, with its history of social justice + public service! 1/
58
19
577
51,451
It was a lot of fun to guest lecture for @OppInsights "Using Big Data to Solve Economic and Social Problems" (opportunityinsights.org/cour…) - it's one of the classes I most wish I could've taken, the students asked great questions, and it's taught in this beautiful building!
1
24
2,309
Spring at Princeton is absolutely stunning :) A pleasure to speak at the Quantitative Social Science Colloquium - many thanks to @PUPolitics for the invitation!
2
61
5,677
Submit your work on applying LLMs for health applications in low-resource languages and settings to the All4Health Workshop! Submission deadline extended to 3/21, workshop website here: nivi.io/all4health.
5
27
4,635
Sunset over snowy @cornell_tech :)
65
3,851
I have an article in the New England Journal of Medicine @NEJM about accuracy and equity in risk prediction, based on my own experiences as a patient. Thanks to colleagues, family, peer reviewers, and editors for very helpful comments. nejm.org/doi/full/10.1056/NE…
1
15
93
18,585
85% of LLM papers studying equity impacts focus on equity *harms* This is a vital discussion. But equally vital is the more opportunity-focused counterpoint: “what new *opportunities* do LLMs enable that could promote equity?” Our new preprint presents 4 opportunities. 1/3
5
28
131
39,216
I am recruiting PhD students for Fall '24! Please apply to Cornell CS if you are interested in ML, data science, health, or inequality. Feel free to retweet! We are based in NYC - here's the view from our island campus (taken during Pride - note the rainbow lights!)
83
883
2,269
469,405
this is interesting!! Relatedly, @_KarenHao, we examined papers about large language models and documented a pronounced US/China schism - American/Chinese institutions collaborated rarely (Microsoft was an exception to this rule). arxiv.org/abs/2307.10700
For decades the US & China have set aside differences to collaborate on science. The partnership has stabilized relations & formed the bedrock of modern research. Now some Republicans are pushing to end it. That could be bad for the US & the world. 🧵 1/ wsj.com/articles/the-u-s-is-…
2
18
123
33,858
Replying to @calimagna
hahaha. yeah, I just checked and apparently other AI researchers have also thought about this quote - nytimes.com/2023/05/01/techn… unfortunately, I don't think the quote actually made it into Nolan's movie, did it?
1
2
277
Sunset landing into Seattle, time lapse.
15
1,254
Finally, this project began as a “paper hackathon”, with 6 of us writing code together for 2 days. It was a fun experiment! We got to know each other (and Trader Joe’s snacks) better (movie theater popcorn = best snack).
1
5
1,819
There are pronounced divides in the collaboration network. American/Chinese institutions collaborate rarely (Microsoft is an exception to this rule); papers with multiple industry affiliations are also rare. 6/N
1
2
22
3,712
LLM authors with different backgrounds focus on different topics: we observe divides by academic/industry, gender, new/experienced LLM authors. 5/N
2
5
1,429
Academic institutions publish more LLM papers than industry institutions, but industry papers and papers with big author teams are more likely to be highly cited. 4/N
1
5
969
The proportion of LLM papers is shooting up: 12x increase in papers using the term “large language model” in 2023 relative to the same period in 2022. The fastest growing topics/keywords on *the entire CS/stat arXiv* are LLM-related. 2/N
1
3
6
1,422
New working paper quantifying arXiv publication patterns in the age of LLMs! Joint work with @rajivmovva, @sidhikab1, @kennylpeng, @gsagostini, and @NikhGarg. We analyze LLM citation patterns, fastest growing topics, many other things. Some of our findings: 1/N
6
32
141
45,814
this is a bit surprising, cuz the COMPAS dataset is super-famous. and indeed, I confirmed that Code Interpreter was aware of its context and the specific finding of disparities in FPR/FNR. and yet it didn't initially perform any of these analyses. 5/N
1
1
498
but its fairness analysis was...very incomplete; it just reported similarish accuracies across groups. (should note - I did not look carefully at its code or check these numbers against previous work, so they may also be wrong) 4/N
1
2
363
but of course the COMPAS dataset is primarily famous for its fairness implications. so I looked at how Code Interpreter dealt with that. It did not suggest any fairness analyses in its first pass, but it did when prompted. 3/N
1
1
388
It immediately ran into some basic analytic issues - label leakage, e.g, producing falsely high performance. 2/N
1
1
408
Boston in the summertime is lovely - so far I have discovered a secret garden by the MIT dome (pictured), a lobster roll salad (very Boston), some excellent udon, and (courtesy of a great recommendation from @dmshanmugam ) a new ice cream place (Toscanini’s - get the B3 flavor!)
2
1
33
3,796
It was an honor to give the graduation keynote at Thomas Jefferson High School for Science and Technology @TJHSST_Official. Congratulations to all the graduates - TJ is a special place, and I will always feel very lucky to have gone there.
1
79
6,476