Has AI been a net positive or negative for survey research?
In episode 1 of The Signal, Dr. David Rothschild (@MSFTResearch, @AAPOR AI task force Co-chair) joins @andrew_j_gordon on where AI helps, where synthetic data breaks down, and why transparency beats standards.
Better models, better RL, better evals, better context - what actually makes agents improve?
We are bringing together researchers and engineers for the next @AgenticAIFdn London event 🇬🇧
Register here:
luma.com/rs92x0u9
Just researchers and technical leaders at frontier labs pondering one of the biggest debates in modern AI.
We hosted our "Who gets to judge the model?" panel with AI Circle inside a cool Midtown NYC apartment, joined by experts from @Amazon and @Google.
Matt Lincoln (@Prolific) made the case that domain expertise alone isn't enough. Evaluators need training in how to evaluate, and rubric design is its own specialized skill.
A big thanks to attendees, moderator Albert Chun, and our panelists who brought real substance to the discussion.
Keep up with our events: luma.com/prolific
Who gets to judge the model?
Prolific and AI Circle are bringing researchers, founders, operators, and technical leaders together in NYC, Sept 16th, for a candid panel on who should actually decide whether a model is good.
RSVP ↓
luma.com/4jliw8f5
🚨 In 24 hours: a fireside chat on what makes a research platform trustworthy with Andrew Gordon and Becca Richards.
We’ve already had great questions come in which we'll be tackling live, tomorrow @ 12 PM EDT / 5PM BST 👉 riverside.com/webinar/regist…
Today we launched The Signal, @Prolific's new podcast on research evidence, not research tools or trends.
David and Andrew discussed what research quality looks like once you remove the human from the process. Give it a watch!
Has AI been a net positive or negative for survey research?
In episode 1 of The Signal, Dr. David Rothschild (@MSFTResearch, @AAPOR AI task force Co-chair) joins @andrew_j_gordon on where AI helps, where synthetic data breaks down, and why transparency beats standards.
Our team recently tested whether AI can replace human survey respondents, using simulated LLM personas against a real US survey (N = 996).
Turns out demographic personas make it worse, idiographic information helps a little, and output format is the biggest lever of all.
Synthetic data predicts human opinion; it's not a substitute for measuring it.
Pre-print: papers.ssrn.com/sol3/papers.…
MTurk's full closure has researchers rethinking online recruitment.
Join @andrew_j_gordon and Becca Richards for a fireside chat on participant vetting, platform design, and data quality, backed by years of independent research.
🗓️ Thu Sept 10, 12:00 PM EDT / 5:00 PM BST.
Great panel with AI Circle in London recently on agent harnesses, evals, and what it really takes to get agents from demo to production.
Thanks to our speakers from @Meta, @ElevenLabs, @OvermindLab, & moderator @josephinePqt1 for leading a rich conversation (and sharing the 📸)
20,000+ doctors, nurses and healthcare professionals across 35+ specialties are now evaluating AI models on Prolific.
Real diagnostic reasoning, real bedside judgment, real safety evaluation.
Every one of them cleared registry verification, specialty confirmation, and a proficiency check built for judging AI reasoning. Less than 1 in 7 applicants get in.
Request a sample dataset: prolific.com/healthcare-ai?u…
🎧 @UMG’s audio wellness venture, Sollos, needed research that could keep pace with a fast-moving team.
Rory Smith and Emma Rugg joined us to discuss why the Sollos team chose Prolific over managed service providers.
piped.video/O341Wq3ddjM