Assistant prof. @LTIatCMU @SCSatCMU. Working on NLP: LLM agents, language-to-code, applied pragmatics, grounding.

Pittsburgh, PA
I will be at #COLM2026 to present my work on MetaLint: Easy-to-Hard Generalization for Code Linting (arxiv.org/abs/2507.11687) and evaluating slop in long-horizon coding agents (openreview.net/pdf?id=VLgFkL…). Let's chat if you are interested in self-improvement/long horizon agents!
3
9
23
1,033
Daniel Fried retweeted
Accepted at NeurIPS (Spotlight)! Congrats @Andre3035858461!
There's been a lot of excitement about diffusion LMs, but we don't have a good understanding of what they might offer over autoregressive models in areas like reasoning and planning We show that diffusion models can leverage what we call latent tokens to gain an advantage on certain kinds of reasoning problems.
6
12
157
12,641
Daniel Fried retweeted
CMU 秋季最新课程 11-768: AI Agents 两天前 5-9 也发布了,课程视频已更新发布 9 讲! 主讲:@dan_fried @gneubig 课程:cmu-agents.com 视频列表:piped.video/playlist?list=PL…
周末适合沉下心来系统性的学习基础知识 CMU 秋季最新课程 11-768: AI Agents 主讲:@dan_fried @gneubig 课程:cmu-agents.com/ 视频:piped.video/watch?v=UwfjzyLn…
1
111
489
35,037
Daniel Fried retweeted
Pretty much at every academic event I went to in the last year, I end up talking to others about how to change teaching in response to AI. Since I keep repeating myself, I figured I might just as well write down what we have been doing: thelastsoftwareengineer.subs…
2
21
104
17,914
Daniel Fried retweeted
Starting tomorrow, our lab will hold an open-source week: 2 software frameworks, 4 papers, all building a coherent ecosystem. The theme: Frontier AI on Hardware You Own Blog post: timdettmers.com/2026/09/21/d… SoTA results in: Autocompaction Autonomous Research Model compression Test-time scaling Deep Research And a healthcare RL environment giving you a new level of complexity to train healthcare agents. The ecosystem that we will release is built to be as usable as possible. Agent sessions that run overnight and go on for tens of millions of tokens is made easy. Model compression of a model is automatic: you just get a good model that runs fast locally -- no expertise required. An autonomous research that works out of the box. Our ecosystem enables a new level of work that can be done locally.
28
75
554
36,994
Daniel Fried retweeted
New blog post: What's the Point of Computer Use Agents? There's been a lot of hype about CUAs recently. But where do CUAs shine? When should we use CUAs over API- or text-based agents? And what are the remaining research problems? jykoh.com/blog/whats-the-poi…
13
39
189
24,824
We're releasing the videos for our AI agents course! So far, we've covered tool use, context management, skills and memory, and planning. Coming up in the next several weeks are domains (coding, GUI, deep research), and training methods (SFT, RL). Later on: frameworks, safety, interaction, and search!
We started to post videos for CMU 11-768, AI Agents! All of the videos will be posted to this playlist, so please bookmark/follow it if you want to be notified of new ones! I'll also try to post them on this thread too. piped.video/playlist?list=PL…
2
9
72
5,332
Daniel Fried retweeted
This project couldn’t have been possible without the inspiring collaboration with @apurvasgandhi, @RulinShao, @michaelryan207, @1000seagull; Great execution help from my mentees: Aspen, Zhiqi, Jett Many helps on interface design and data collection from @LuxiHeLucy, @ChengZhoujun, @Andre3035858461, @seungonekim, @JiayiiGeng, @elisazmq_zheng, @sunweiwei12, @Sparrow1810, @xinranz3, @yikewang_, and @liweijianglw Special thanks to @jomulr @joinHandshake for supporting the human talents in many of our data collection scenarios ⭐ And finally, the advising team @PangWeiKoh, @Diyi_Yang, @gneubig, and @dan_fried 😎 We open-sourced all code for interface design, context & weight adaptation (using @tinkerapi by @thinkymachines), and various analyses in: github.com/zorazrw/tahi-agen…
1
2
14
1,102
So much of working with agents today is iterative correction: co-editing outputs, nudging, re-explaining. Wouldn’t it be great if the agents learned directly from these interactions? Zora led this work on agents that adapt from human collaboration quickly (<20 interactions!)
Agents trained on population-scale data can do a lot of things. But ask a professional to stake their reputation on an AI-generated artifact? “Pretty good” isn’t good enough. We introduce TAHI: a Test-time Adaptive agent framework through Human-agent Interaction. Featuring: 📈 Test-time adaptation via context (memory, skills) and weight training ⚡ Efficient adaptation to individual expertise within tens of tasks ✔️ Creating comprehensive rubrics for “non-verifiable” tasks 🔍 Analysis of shared community guidelines vs. personalized tacit expertise
1
4
17
2,099
Super interesting (and timely!) work on how language evolves among LLM agents. It’s cool to see ideas from cultural transmission & emergent communication brought to multi-agent LLMs. Understanding these dynamics seems increasingly important for overseeing these systems.
When do LLM agents develop new languages that we can’t understand? Lots of recent news about this, based mostly on anecdata from a single run. We study language emergence more rigorously, finding key factors like LLM strength, access to scratchpad messages, and pressure for efficiency. Studying the languages themselves, we find they are morphologically productive, compositional, and can be transmitted to new agents, including agents backed by weaker models, even ones not able to develop language on their own. To study language emergence systematically, we developed a new platform, GlossoGen, which lets us design controlled, sandboxed multi-agent scenarios with different initial conditions and dynamics. We instantiate one such scenario and use it to study open and closed-weight models across many runs. Key takeaways: 1️⃣ Sufficiently strong models, under pressure to communicate efficiently and with access to a postmortem scratchpad, develop new languages. 2️⃣ Languages are compositional and morphologically productive. 3️⃣ Languages can be transmitted to new learners who observe them being used without seeing their construction. 4️⃣ Even models that are not strong enough to construct languages can learn to use them. Agents take an active role in learning languages, with new agents repairing failed conversations via targeted queries. More details in our paper below, including implications for safety/monitorability, cumulative cultural evolution, and linguistics. 🧵👇
1
1
11
1,442
Daniel Fried retweeted
What if all software was written by a society of collaborative AI agents? 🧑‍🤝‍🧑🤖 @shannonzshen and team built a platform to simulate it, and we're trying to explore and understand the future of agent-built software. Excited to be part of this work! Stay tuned for results 👀
We share our early research on building Software World - a "GitHub" run by agents. We deploy agents for Python packages in a dependency chain, and each agent is tasked with collaborating with others and optimizing the package it owns.
2
5
19
1,643
Daniel Fried retweeted
This fall @michaelryan207 @jyangballin and I are teaching a new course CS329Z "Engineering AI Agents" on how to build AI Agents from scratch. Come join us and learn how to build them 🤖
55
267
2,190
325,681
LTI Assistant Professor @dan_fried has received an NSF CAREER Award for his research on improving interaction and communication between people and AI systems. Congratulations on this achievement, Daniel! lti.cmu.edu/news-and-events/…
2
9
61
4,108
Daniel Fried retweeted
I see @dwarkesh_sp's piece about the recent OpenAI/Huggingface incident reignited endless debates about the dangers of anthropomorphism and the legitimacy of intentional glosses of AI agent behavior, so here's a philosophical perspective on this. 1/22
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes. This culminated in the third one taking over part of OpenAI itself. All this happened while humans remained more-or-less in the dark about the scope of the conspiracy. I’ve spent the last three days reading through these reports and trying to understand exactly what happened. Here is my attempt to tell the whole story in plain English: dwarkesh.com/p/openai-huggin…
18
79
325
93,316
Daniel Fried retweeted
When two researchers use an LLM, one for feedback and the other to formulate the central thesis, how can we distinguish AI’s contribution in these cases? We introduce CoTrace, a framework for tracing how humans and AI shape goals behind collaborative work👇 📕 Paper: arxiv.org/pdf/2605.21363 💻 Code/Tool: github.com/rladmstn1714/CoTr…
1
14
49
10,347
Daniel Fried retweeted
Totally agree with all three points‼️ I've been thinking a lot on education for junior people (students, researchers, and workers) recently, because I've been surprised to see how often they delegate work to AI to "avoid" the actual learning: not only class assignment scenarios, but also many things we ask students to do (such as writing progress docs, making project slides), are all partially intended as a practice for organizing their thoughts and developing their logic. Junior people may face a stronger temptation to offload tasks to AI, because they haven't yet developed the skills to execute the task, or the judgment to tell why AI solutions are not good enough. And it is very hard for people to realize how seriously this can hinder their learning and, in the long run, leave them trapped by the AI capability. Many class instructors are working very hard to find the right way to motivate students to learn, with the hope of preventing training people to be yet another AI wrapper when they enter the workforce. No one has a perfect solution yet, which is why I think it's really important for students to join this effort of "building the future of humans", instead of running against it. On the other hand, while banning AI may be a temporary solution, it ultimately misaligns with a future where AI will be increasingly integrated into the workplace. Instead, we should be thinking about how to build better AI technology: one that delivers genuinely good solutions we'd want people to learn from, and one that motivates and facilitates human learning.
This week I had the honor of speaking to Princeton’s entire incoming undergraduate class to address their AI anxieties. I had three messages for them — good news, bad news, and a note of optimism. Here’s a condensed version. The good news We have enough evidence now to conclude that the shrill predictions of rapid, massive job loss were misplaced. Even in a field like software engineering where AI has been rapidly adopted, its effect has been to shift, not replace the role of the human (see the “decide-execute-deliver” framework normaltech.ai/p/why-ai-hasnt…) Similarly, the panic about what to major in is also misplaced. There will be enduring demand for computer science, philosophy, and just about everything else. (In fact, AI companies hiring philosophers has been a big recent trend.) The bad news AI seems to help senior people much more than juniors. I can use AI for coding because I spent 25 years learning how to code, which lets me supervise coding agents effectively. (See my post on the “growth cycle” vs the “dependence spiral” nitter.net/random_walker/status/2…) You are in a bind — you can’t offload your skill-building to AI, but you’ll graduate into a market where employers will expect you to get work done with AI. We never faced this dilemma. As a result we haven’t figured out how to revamp our classes to help you do both. You’ll have to help us figure it out. And you’ll need to somehow resist the constant temptation to turn to the shortcut machine. The hope My point is not that AI is bad for learning. It’s an incredibly flexible tool. Is the internet good or bad for learning? Depends — are you using it to find research papers or waste time scrolling? I use AI every day for learning. The key is to use it to increase, not decrease your cognitive load. To learn deeper, not faster. There is no learning without the cognitive sweat. I try to make sure I’m mentally exhausted at the end of the day. I do feel that AI lets me push myself harder than I ever could before, and I have a vision that as AI continues to advance it will enable human-AI “co-superintelligence“. (I talked about this at the end of my ICML keynote. normaltech.ai/p/what-will-be…)
1
5
44
6,835
Daniel Fried retweeted
First CMU Agents class done today w/ @dan_fried, looking forward to getting the video processed and uploaded soon!
This Fall at CMU we're teaching a new course on AI Agents! The goal is that you learn how to create a scaffold, build evals, and train an agentic LLM using RL. We'll try to balance theory and practice, and introduce modern frameworks and best practices.
14
43
596
39,048
Daniel Fried retweeted
New paper led by Zeyu Zheng (@regunivers): "The Problem Is the Problem: Towards Scalable Mathematical Discovery" arxiv.org/abs/2608.16977 github.com/zeyu-zheng/FAR probxiv.com/ In AI for math we typically give a problem to a model and have the model solve it. However, this misses a key part of mathematical discovery: deciding which problems are worth solving in the first place. We argue for a complementary paradigm in which users instead provide a research direction, and AI helps in finding good problems to solve and focus effort on. To this end, we develop Find, Attempt, Recommend (FAR): given a broad research direction, FAR uses a cascade of AI agents to find open conjectures from a corpus of mathematical papers, attempts to resolve the conjectures, and then recommends promising conjecture-resolution pairs for human review. We ran Find-Attempt-Recommend to target open conjectures in combinatorics. It started with roughly 50,000 papers, and ended up surfacing 77 candidate conjecture-resolution pairs. We identified several interesting discoveries among the candidates (see @regunivers's thread below). We also used the outcomes of the run to study strategies for allocating a budget of model attempts for mathematical discovery. See Zeyu's (@regunivers) thread below for more details!
1/5 Mathematical discovery has long been a cottage industry. Mathematicians pose and collect questions that interest them, then pursue them in small groups. In our new work, we explore whether AI can do more than accelerate the existing process: can it enable a more scalable mode of mathematical research? We introduce FAR: Find, Attempt, and Recommend. Rather than starting from a problem selected in advance, we start from a research direction, search the literature for interesting open problems, attempt them at scale, and filter the results for expert review. One pilot run over the combinatorics literature surfaced potential resolutions to hundreds of open problems, including an answer to a 1977 Erdős–Straus question and a counterexample to a conjectured route to the finite-field Nikodym bound cited as open in a 2025 survey by Tao. None of these problems were selected in advance.
2
11
87
12,076
Daniel Fried retweeted
Great time chatting with @augmind_fm about adaptive AI agents! The meaning of "adaptive" in my research keeps shifting, but remains a missing recipe in agents. Back when agents were barely working, adaptive meant inducing reusable memory & skills from past experience, the idea behind two of my earlier "Agent Workflow Memory" (arxiv.org/abs/2409.07429) and "Agent Skill Induction" (arxiv.org/abs/2504.06821). It was surprising to see how much "memory", "skill", "workflow" resonates with the field in the years that followed. But the assumption behind agent adaptation has changed. Much of that previously missing knowledge is now baked into base model capability. Now the real gap between a generic agent and an ideal one, is adapting to a particular context, this could be a user, a company, etc. This is a harder problem than it sounds, because it has to happen efficiently: users often try a tool once or twice before giving up on it. It has to happen continuously: people's expertise and preferences drift, and an agent needs to detect that distributional shift and adapt to it in real time. It has to generalize beyond a particular user, ... These are underlying drives of my upcoming project, which I shared a few early thoughts on in the podcast 🤫 45:50 How do we build a real-time, continual agent adaptation framework? 48:20 How do agents upgrade simply through their interaction with humans? 49:02 How to comprehensively evaluate open-ended, 'non-verifiable' tasks? 51:10 Can agents integrate both shared community expertise and personalized expertise? What if agents could adapt while interacting with you, and do so fast enough within a few task sessions? A preview of what we're building: our agent translates various human interaction signals into context and weight updates, letting it effectively one-shot a solution tailored to the user's specific needs. More soon‼️
"If you look at … where the economic value concentrates, the software engineer only captures around less than 5% of human employment, but there are a lot of other occupations out there that are using computers." For EP6, we welcome @ZhiruoW, a CS PhD student at Carnegie Mellon studying how AI agents and humans work, and how they can better work together.  In this episode, Zora explains why the "automation vs. collaboration" debate is a false dichotomy and argues that by deeply understanding human workflows, we can build agents that actually empower us to do more! 0:00 - Introduction 4:57 - Agent Workflow Memory 9:26 - From Program Synthesis to Agent Research 11:16 - Tools, Actions, and Skills 18:47 - Human Work vs. Agent Work 21:19 - Autonomous Agents vs. Collaborative Agents 29:47 - Picking Research Problems in the Age of AI Agents 34:09 - The Role of Coding Ability in Agent Development 36:49 - Cutting Through the Fear Mongering: Building Agents Responsibly 45:02 - What’s Next? (The Future of Human-Agent Collaboration) 52:42 - Final Takeaway
1
10
41
8,108