Research Scientist at @SFResearch

Palo Alto
InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation Paper: bit.ly/3P4cQj9 How well can LLMs simulate real people? Most evaluations rely on surveys or questionnaires as proxies — never checking against what individuals actually said. 🔍 InterviewSim introduces an interview-grounded evaluation framework at scale: 671K+ Q&A pairs extracted from 23K verified interview transcripts across 1,000 public personalities, averaging 11.5 hours of content each. 📊 The framework evaluates simulation fidelity across four dimensions: → Content similarity → Factual consistency → Personality alignment (Big Five) → Factual knowledge retention (MCQ) 📌 Key finding: grounding in real interview data substantially outperforms biographical profiles or parametric knowledge alone. But how that data is used matters — retrieval-augmented methods capture personality style best, while chronological methods better preserve factual consistency and knowledge retention. 💡 The work also reveals that question type is a stronger predictor of difficulty than method choice. Social identity questions (birth dates, family details) yield the highest contradiction rates across all methods, while motivations and values questions are most forgiving. Authors: Yu Li @yooli23, Pranav Narayanan Venkit @PranavVenkit, Yada Pruksachatkun @yadapruksachatk, Chien-Sheng Wu @jasonwu0731 #FutureOfAI #EnterpriseAI #NLProc
1
7
34
2,884
Demographics aren't enough to simulate human behavior. 🧠 New research introduces SCOPE: a framework and persona dataset collection that moves beyond demographic templates to build richer AI personas grounded in sociopsychological structure. 📄 Paper: bit.ly/49ZduoK Key findings across 7 models: Demographics alone explain only ~1.5% of variance in human responses | Adding traits, values & identity narratives improves behavioral alignment while reducing bias | SCOPE personas outperform existing approaches on SimBench, an external social and behavioural benchmark. The work also shows SCOPE can augment existing persona systems like NVIDIA Nemotron to improve performance. ✨

Our key finding is that existing persona frameworks are often superficial or overly broad. A more socially sensitive approach produces personas that are less biased and more representative. Synthetic persona creation pipeline and corpus now available for research use. 🔓 Authors: Pranav Narayanan Venkit @PranavVenkit, Yu Li @yooli23, Yada Pruksachatkun @yadapruksachatk, Chien-Sheng Wu @jasonwu0731 #FutureOfAI #EnterpriseAI #AIResearch #AgenticAI #ResponsibleAI
3
11
1,212
Yu Li retweeted
To effectively solve modern computer tasks, AI agents need to be able to strategically explore the environment and efficiently learn from past interactions. We present R-MCTS and Exploratory Learning for building o1-like models for agentic applications. Our GPT-4o powered R-MCTS agent creates SOTA performance on VisualWebArena. Notably, R-MCTS and Exploratory Learning (without MCTS) demonstrate the compute scaling properties in both training and testing time! 🌐: agent-e3.github.io/rmcts-exp…
1
7
19
6,642
The 6th Workshop on NLP for ConvAI at ACL 2024 will take place on August 16th in Thailand! 🌟 Join us for an inspiring day filled with insightful talks and groundbreaking research. Check out our schedule and meet our amazing speakers! sites.google.com/view/6thnlp… #NLP #ACL2024 #AI
2
6
559
Today we officially release ✨Vision-Flan✨, the largest human-annotated visual-instruction tuning dataset with 💥200+💥 diverse tasks. 🚩Our dataset is available on Huggingface huggingface.co/datasets/Visi… 🚀 For more details, please refer to our blog vision-flan.github.io/index.…
1
16
47
6,557
📣 Happening now at #ACL2023, #NLP4ConvAI workshop in Harbour B! Thrilled to have the brilliant Diyi Yang as our first invited speaker. 🗣️ She's diving into "Inclusive Conversational AI for Positive Impact" - a must for everyone in #NLP 👏💡 #ConversationalAI
1
3
11
4,153
🔔 Next up at #NLP4ConvAI #ACL2023NLP , we're thrilled to welcome Nurul Lubis from Heinrich Heine University Düsseldorf! Get ready to explore "Dialogue Evaluation via Offline Reinforcement Learning and Emotion Prediction" in Harbour B. Don't miss it!
1
2
310
🔔 Excited to have Jason Weston from Meta AI as our next speaker at #NLP4ConvAI #ACL2023. He'll delve into “Improving Open Language Models by Learning from Organic Interactions”. Join us in Harbour B for an insightful look!
1
1
295
Replying to @yooli23
Congrats!
1
1
131
#ACL2023NLP #NLP4ConvAI @5thnlp4convai Our workshop starts with @Diyi_Yang and @larry_heck's invited talks!
5
33
9,208
🚨 LAST CALL for #NLP4ConvAI Workshop direct submissions! 🚨 Don't miss this opportunity to showcase your innovative research for Conversational AI. Submit your papers TODAY to join the brightest minds in the field! 🔗 sites.google.com/view/5thnlp… #NLP #ConvAI #ACL2023
2
2
627
Thinking about doing a PhD in Computer Science? Sign up for Columbia’s Pre-Submission Application Review program! Through PAR, a current PhD student who works in your research area will review your personal statement and CV. Apply here: docs.google.com/forms/d/e/1F…
2
2
10
Yu Li retweeted
How do you evaluate your dialog systems?🤔 Our lab's LegoEVAL is here to help. It is a tool for dialog systems evaluation. @aclmeeting demo track. #Preprint: arxiv.org/pdf/2105.01992.pdf
4
11
78