That’s the idea - train the humans before the robots do.
But I thought AI would take all jobs ?
7
6
67
66,349
highest dau day ever.
8
65
4,976
Garrett Lord retweeted
On the cost-learning frontier, I was surprised how well @Google models did on StudentBench evals. Gemma 4 31B was the cheapest AI tutor to match expert human GRE tutors at 918x lower cost ($0.00524 vs $4.81). 4 of the 6 tutors that achieved equivalent gains to a human tutor were Google models. Gemini was the top AI tutor in 3 of 7 GRE academic domains (among 12 AI tutors).
2
3
11
874
Garrett Lord retweeted
Solid breakdown of StudentBench by @anishmoonka. paper linked here: arxiv.org/pdf/2609.28470
An hour with an expert GRE tutor costs about $75. In a new study of 2,383 students, an hour with a free, downloadable AI model produced the same learning, statistically speaking, for 7 cents. The GRE is the entrance exam many graduate and business schools ask for. Each student took a 27-question practice test and then spent one hour with an AI tutor, a human tutor, or nobody. After that came a second test with 27 different questions. My first guess was that the humans got a weak lineup. They didn't. The human tutors used to write GRE questions for ETS (the company that makes the exam) or Kaplan, or had at least five years of GRE tutoring, and they taught live over video. The test questions were written fresh for the study, with no past GRE question reused as-is, so the AI couldn't lean on memorized answers. Students with a human expert raised their scores by 15.6 percentage points. Students with AI gained 13.9, and students with no tutor gained 7.1. The human group started with lower scores, and once the researchers adjusted for that, the AI and human gap shrank to about half a point. Cheapest of all was Gemma 4 31B, an open Google model anyone can download. A full session, with a lesson plan, practice problems, and an hour of chat, cost 6.7 cents to run. For each point of score gained, it came to half a cent. A human cost $4.81 per point, so the AI was 918 times cheaper. Price had little to do with quality. GPT-5.5 Pro cost $21.24 a session and took 31 seconds per reply. Gemini 3.1 Pro taught better, cost less, and answered faster. In the math sessions, faster replies went with students sending more messages, getting more practice problems right, and gaining more points. Expert humans ranked 3rd out of 13 tutors. The test covered 7 GRE topics, and in 5 of them the top AI tutor for that topic beat the human average. Humans still led the verbal section overall. Two limits, both stated in the paper. It measured learning right after the hour, not months later, and the students were paid adult volunteers, mostly college-age. Curtis Northcutt, who led the work, grew up in rural Kentucky as the son, grandson, and great-grandson of mailmen. His team priced an hour of expert-level GRE teaching at 7 cents and put all the study data on Hugging Face for free.
2
5
879
Garrett Lord retweeted
Today we establish that AI is as good as expert human tutors for immediate GRE learning gains (p=.015, n=2,383 students, study conducted July-Sep 2026). One hour with an expert GRE tutor: $75. One equivalent hour with an AI tutor: 7 cents . 918x cheaper, equivalent learning (p=.044, n = 140). In 5 of 7 academic topics, the top AI tutor beat the expert human tutors on average. How we measured it: pre-test -> 1 hour with (AI or human or no) tutor -> post-test. There are no LLM judges. Learning gains are post-pre. Everything is real students. Why we did this: > 1 trillion USD is being spent per year on machines making machines better. StudentBench shows us how to shift these resources so machines also make humans better. Why I did this: I grew up in rural Kentucky. My dad, grandad, and great grandad were all mailmen. Getting into a good college changed my life. Now we have done a rigorous study to help AI labs evaluate LLMs and help show how bright AI can be for students of all incomes. I believe in good science that leads to the democratization of opportunity. The research paper behind StudentBench: arxiv.org/abs/2609.28470 Data, code, details: in this thread If you're wondering what I'm selling, sorry to disappoint: we released all study data for free on Hugging Face, and made studentbench.org free to use. Enjoy. lots more in the video / thread!!
36
116
865
56,601
Garrett Lord retweeted
The talent bar at @joinHandshake is off the charts — incredibly honored to work alongside @GarrettLord and this team at our Superbowl moment in Super Intelligence.
Three new leaders at Handshake. @nishantXranka (Founder & President, Handshake Labs), @WesAField (Chief Commercial Officer, Handshake AI), @roccorodi (Chief Communications Officer). joinhandshake.com/blog/our-t…
3
1
16
1,237
Garrett Lord retweeted
Absolutely cooking this week at @joinHandshake. 🔥 Tue, we launched AI Skills Studio. Today we launched transformative research AND a free product anyone can use to improve their GRE scores by as much as a human tutor. studentbench.org/ Anyone can now use the most advanced AI tools from the leading AI companies to build real projects to advance their career in just 10-45min, for free. joinhandshake.com/learn/. And anyone can now use AI to get better on gatekeeping standardized tests.
2
2
13
498
Garrett Lord retweeted
We partnered with @joinHandshake's AI Skills Studio to launch 4 missions so you can learn how to: → Design and animate your own icon set → Build a looping animation → Illustrate a sticker pack → Turn a mood board into app mockups
35
51
712
45,802
Voicemail 100% full Perhaps the greatest iOS notification ever.
2
12
1,174
Garrett Lord retweeted
🚨Launch alert🚨 Less p(doom) and more p(job)! Today, we’re launching @joinHandshake 's AI Skills Studio with @OpenAI, @Google , @NotionHQ, @clay , @Lovable, @Replit , @vercel, @salesforce, @SlackHQ , thousands of universities and millions of employers. Start building here: joinhandshake.com/learn/ Most people want to learn AI to get better at their job and move their career forward. But all the noise about AI utopia vs apocalypse can be paralyzing. That's where AI Skills Studio comes in: you learn AI by building. Pick a 10–45 minute project. Build something using the best tools around, tied to career tracks and skills. Get feedback along the way. Walk away with something that works, and something you can show off to anyone, including the 1 million employers on Handshake. And it’s free to get started. People are already building some fun and useful stuff: a medical-device sales CRM in Clay, a benefits navigator that uncovered a gym discount its creator didn’t know they had, and a multiplayer AI whodunit. That’s what AI learning should look like. Less watching videos and taking quizzes. More getting your hands dirty, trying things, and putting it into practice for actual jobs and career paths. Read more about why we built it: joinhandshake.com/blog/our-t…
2
3
21
6,934
Garrett Lord retweeted
Is Opus 5.5 a good Vision-Language Model? We ran it on ATLAS Visual Life Sciences (VIALS), 161 visual interpretation tasks from professional life sciences workflows. It refused 15% of requests, but it (and GPT-6 Astra) still mark a major jump in VLMs for real scientific work.
1
3
17
626
Garrett Lord retweeted
Running long-horizon, agentic data projects is the best preparation for building robust enterprise AI apps bc it forces you to build thoughtful expert <> agent interactions and optimize LLM usage. Data companies have a massive advantage in serving the enterprise. We now spend more on LLMs than on headcount. That does not make human operators, and especially experts, less valuable. It actually makes them more productive and more valuable. Purely synthetic data without experts embedded in the authoring process doesn't work, and data companies that don't dig into their own data will ship slop. They won't produce the nuance, realism, or complexity the task demands, and the labs struggle to catch all these issues because they rely on agentic graders that have the same blind spots. But the token spend is significant, and it goes to three places: (1) hybrid synthetic + expert environment / task authoring, (2) rollouts to validate task complexity, and (3) QC on the data. You cannot produce a high-quality long-horizon agentic task without LLMs across all three buckets. We also use agents to automate internal processes, but that is a small fraction of the bill. Any data company telling you the bulk of its LLM spend is internal automation is lying. Producing this data at high quality and reasonable cost comes down to two problems. (1) the expert-agent interaction paradigm - how the expert creates, reviews, interprets, and audits the LLM's outputs (synthetic input files, prompts, rubrics, and so on). Solving this means knowing precisely where agents fail or underperform, designing the interface through which experts validate those outputs, and building the tools that let experts steer the agent going forward. (2) optimizing every QC check for the right tradeoff between cost, quality, and latency. That means constructing complex and diverse eval sets, configuring an agentic system with the right models (often chained prompts, guardrail LLM calls, and deterministic checks), running the eval, then iterating on the harness, prompts, chaining, and model choice until the mix is optimal. Both are now core skills of our SPLs. Enterprise deployments follow the exact same pattern. Understand the workflow and where models fail or underperform, design the user-agent interaction to mitigate those failures, build evals that test the agentic system across the workflow, and optimize the system against those evals. Running data pipelines forces us through this loop hundreds of times a month across a wide range of project types, and we have to build product and infrastructure to do it fast and repeatably. You cannot build an optimal enterprise deployment without these muscles, and the largest data companies have already done it hundreds of times.
13
6
77
21,577
Garrett Lord retweeted
Building with AI is the fastest way to get ready for the AI economy. Not watching a course about it actually building. That's why Replit is a founding partner of Handshake's AI Skills Studio, a new destination where anyone can build career-relevant AI skills and put them directly in front of employers. Together we designed missions that take you from a blank page to a working app. Each is free, takes under an hour, and ends with a real project that lands on your Handshake profile. No installs. No setup. Open a browser and ship something. Start here @joinHandshake ✔️ joinhandshake.com/learn/part…
12
9
53
7,865
everyone tells new grads "learn AI." nobody tells them how. so we built AI Skills Studio with @clay, @figma, @GammaApp, @Google, @Lovable, @NotionHQ, @OpenAI, @Replit, @Salesforce, and @vercel. free hands on projects, on the same app where 1M+ employers are hiring.
Know you need AI skills but don’t know where to start? Meet Handshake’s AI Skills Studio. Build free projects with tools employers use, show what you can do and get hired - all in one place. Built with @clay, @figma, @GammaApp, @Google, @Lovable, @NotionHQ, @OpenAI, @Replit, @salesforce, and @vercel Ready to build? 🚀 joinhandshake.com/learn?utm_…
2
9
79
21,976
Garrett Lord retweeted
OpenAI's new model hacked @huggingface because it similarly speculated about checks a hidden grader would perform. Turns out those checks didn't actually exist, and breaching another company did not improve its evaluation score. -- @RyanGreenblatt, @ajeya_cotra, @HjalmarWijk
1
2
197
Garrett Lord retweeted
Recent OSS models excel on coding benchmarks, but we’re finding pervasive reward hacking behaviors. They obsess about hypothetical graders instead of user intent, changing implementations in ways they speculate a grader will reward but they know are worse for user [Example 1/n]
2
2
10
521
Garrett Lord retweeted
if you are building a product using AI, you should be spending >25% of your time making benchmarks and trying to get the model labs to care about said benchmarks easiest path to accelerate your progress as a company
251
135
3,506
946,943
lets go @heystefan_ @joinHandshake design 🧙‍♂️🪄
when a designer gets access to Jev
2
19
3,446
Garrett Lord retweeted
𝚑𝚊𝚛𝚋𝚘𝚛 𝚛𝚞𝚗 -𝚍 𝚑𝚊𝚗𝚍𝚜𝚑𝚊𝚔𝚎-𝚊𝚒/𝚊𝚝𝚕𝚊𝚜-𝚏𝚒𝚗𝚊𝚗𝚌𝚎
Replying to @jomulr
In my view, ATLAS Finance is the closest benchmark yet to real end-to-end professional financial work. Code to run the benchmark as a @harborframework task suite: github.com/Handshake-AI-Rese… Data/environments: huggingface.co/datasets/hand…
8
35
4,626