intelligence farmer. member of semi technical staff. prev @mercor_ai @scale_ai @McKinsey

San Francisco, CA
it's only called human data if a human clicked submit. otherwise it's just sparkling synthetic data
3
3
51
5,720
one of the most important new benchmarks far more relevant to most people than PhD-level reasoning or olympiad math
📣 Excited to introduce OfficeQA Pro V2, the next generation of OfficeQA Pro! It's built on a new corpus of 120,000 PDFs provided by the U.S. Treasury. Frontier AI agents average 26% accuracy. More info below 👇 databricks.com/blog/introduc…
1
10
2,453
Sid Potdar retweeted
Of course that’s your contention. You’re just getting into an RL data startup. You just got finished reading some thread about how the harness is better and the scaffolding is what makes the difference. You’re gonna be convinced of that til next month when you decide the model is what actually matters, and then you’re gonna be talking about pure capability and how harness doesn't mean much. That’s gonna last until you read more tweets and say the old model vs harness dichotomy is already obsolete and the RL-harness lifecycle with pure outcome rewards fail on long-horizon non-verifiable work because of credit assignment then talk about process supervision without knowing what it even means. Then you're gonna read some technical blogpost explained to you by sol as ELI5 about how cheap reliable process-level signal+solving credit assignment at agent scale is what actually matters. and in the end you'll realize that you've been following what everyone is coming up with no original thoughts while burning VC money because you can't even do B2B sales properly when you could've just done powerpoints made by LLMs and shilled it on linkedin
23
20
444
25,412
increasingly suspicious of long-horizon tasks being a psyop by Big Token to prop up API revenues nothing I’ve ever accomplished has required more than one context window
1
10
1,389
Sid Potdar retweeted
the extent to which AI progress is data bound is under appreciated by most
172
69
1,552
222,065
amateur hour behavior. you only have to slice the salami that thin if you don’t have much meat to start with
Salami Slicing In a transactional context, sometimes your counterparty may try to essentially scam you for a small margin -- so thin that they hope you do not notice, or let it slide because it basically doesn't matter. By doing this over and over again, your counterparty hopes to accumulate some meaningful aggregate advantage. This is an amateur technique, because Salami Slicers usually do not realize that: 1. People notice. They may not say anything. But oh man do they notice. 2. Being recognized as a Salami Slicer is extremely negative and will kill you in social and business contexts, because you are showing that you will compromise your ethics for trivial amounts of money (or whatever else is at stake). Nobody will trust you for anything significant if they know you're the type of degenerate to steal pennies while others aren't looking. In fact, frequently, people will permit others to Salami Slice them as a kind of test -- will you do the wrong thing, thinking you can get away with it, for a small amount? A valuable signal. Thank you. Skilled Salami Slicers will maintain some level of plausible deniability -- "I'm sorry, I forgot to keep this minor obligation", "oh I wasn't sure whether you were paying this or I was", "I thought I didn't have to do this thing because of that other thing", etc. But there are patterns to this, and you can sniff it out quickly. (For example, Uber's ETAs are always optimistic -- the cars are always late or on time, but never early. Huh.) Some people will go 0-to-100 when they notice they're being Salami Sliced, because they understand what it says about their counterparty's view of themselves: it's extremely dehumanizing. If you try to strategically and deliberately deceive me for minuscule gains, then clearly your view of me is so low that you would cause me unlimited suffering if it was to your benefit. In a Schmittian sense, it's a hardcore Friend-Enemy distinction. Amateur Salami Slicers are occasionally surprised when they are spotted and receive the full Fuck You I'll Kill You, scorched-earth zero-mercy response: their counterparty understands the full symbolic reality of Salami Slicing. I hate being Salami Sliced.
1
5
1,157
Sid Potdar retweeted
Today, we remember a legend. On this day in history, Harambe would have celebrated another birthday. An icon that became part of internet history, American culture, and an entire generation’s timeline. Tomorrow marks 10 years since we lost him. Ten years since the moment the world stopped scrolling and collectively mourned something bigger than a meme. He became a symbol of loyalty, strength, chaos, unity, and the strange beauty of the internet bringing millions of people together for one cause: never forgetting Harambe. Everyone remembers where they were when they heard the news. And somehow, a decade later, his legacy still lives on. Gone, but never forgotten. Rest easy to a true patriot. 🕊️🇺🇸 May 27, 1999 — May 28, 2016 Forever in our hearts.
6,424
20,212
151,201
24,263,836
saw trajectory’s great work with mercor up close, excited for them to make continual learning possible for every org
Today, @MichaelElabd, @QuantumArjun, and I are excited to announce Trajectory. We are a research lab and product company building the platform for Continual Learning. Our platform unlocks the signal already sitting in product usage, so companies can continuously post-train large-scale agentic models that outperform the frontier. @trajectorylabs We’ve raised $15M from @Conviction, @BessemerVP, @radicalvcfund, @jeffdean, @drfeifei and more. We’re partnering with some of the best AI-native companies: @ClayRunHQ @Harvey, @DecagonAI, @mercor, @RogoAI to power their agentic systems, some of which we are already in production with. We’ve brought together a world class research team from DeepMind, OpenAI, Apple, Meta Superintelligence, Amazon AGI, Scale AI, and an elite product team from Stripe and Figma. AI will never again start on day one. Every correction, every retry, every edit will make products smarter. This is Continual Learning.
1
11
2,136
Sid Potdar retweeted
I’ve left Google DeepMind. The last two years have been an incredible whirlwind. A couple years ago, I joined a small startup called Codeium. There, I got to ship Windsurf, train SWE-1 (a frontier agentic coding model), go to DeepMind in the $2.4B acquisition. Now, I decided to leave the acquisition money and DeepMind. I’m grateful to the mentors, teammates, and friends I worked with along the way. At Windsurf, thanks to @_mohansolo and Douglas Chen, I got to see what a fast moving startup that ships relentlessly and builds for the future looks like. I learned from @thenickmoy how excellent research leadership can drive outsized innovation. At DeepMind, I got to push the frontier of agentic coding, be part of the amazing team that shipped Antigravity and contributed to Gemini 3. DeepMind is a rare place: deeply curious people, exceptional research taste, and access to enormous compute and Google-scale infrastructure. A few things that I learned: 1. Finding the right hill to climb. Now more than ever, there are a multitude of directions to push the frontier in AI research. It’s easy to optimize for the wrong benchmark or capability. You should step back regularly to question if you are climbing the right hill, and adjust course often. 2. The secret to being a fast-moving team. Moving quickly is not just about working hard and long hours. It requires making concrete bets about where the world will be in 6 months, aligning around them, and cutting everything else. This was our journey from the Codeium Extension → Windsurf IDE → SWE-1 → Antigravity → Antigravity CLI 3. Silicon Valley is small. Since the split of Windsurf to DeepMind and Cognition, many of my colleagues have gone to other exciting places - Thinking Machines, OpenAI, xAI, Cursor, fast-moving startups, or started their own companies. I’m grateful to have worked with so many talented, hungry people whose stories are not yet finished. So what’s next? We are living in one of the most exciting and powerful times in human history. Just like we transformed software engineering, soon every industry, every unit of work will be radically transformed, democratized, accelerated. With this comes new challenges, and new doors of frontier research to be opened. More soon.
114
34
971
297,075
Sid Potdar retweeted
If you think you are above labeling data you are not going to make it
16
14
207
23,001
Sid Potdar retweeted
Solving the science of asset selection in a future (or indeed the present) where every company is a "Context Acquisition Company" is the real frontier. I love that everyone is getting around to the idea that the secrets (scarce context) currently illegible to/hidden from computers (human or machine) are everything. Now the next leap for people to make is that the science of sourcing, selecting, and monopolizing that context (really THE ASSETS that produce it) is everything. If AI progress is a function of compute and data (most algorithmic progress is really just data progress; h/t @BerenMillidge, @_kevinlu, @mentalgeorge, @GarrettLord, etc.), then every company is going to have a context desk just like they will (or already do) have a compute desk. The difference is, CONTEXT IS NOT FUNGIBLE. Most context (both that exists right now and that will be created in the future) will be completely commodity beta. Winning will be about getting to and instrumenting the right asset (context production factory) first. And yes, there are right and wrong answers. To do this kind of asset selection well requires an extremely scarce meta-capability: the ability to coordinate the right kind of access and the right kind capital at the right time. These assets (and the secrets within them) are structurally difficult to access, evaluate and instrument. They are not floating around in banked processes, to be frictionlessly purchased on listed exchanges, or willingly coming through Mercor or Handshake's expert portal. (Yes, a context production asset can be (very often is) a single person or collection of people.) When @WillManidis talks about a Deal Guy Yuga, what he means is that there are people who have deeply internalized the fact that at the limit, in a world of infinite intelligence, access to/monopoly on the right permissioned data streams is all that matters. Getting yourself to a position (meta-access, meta-capital) where you have the ROFR on those permissioned data streams, means being a generational Deal Guy. This is a very different and specific kind of "Deal Guy" though. Knowing which asset(s) are going to give you the right context to create, compound, and commercialize the best vertical world model now and into the future is the new form of security analysis. But the triple-exceptional combination of domain expertise, meta-access, and technical ability that’s required to execute this new security analysis effectively is scarcer than the talent at quant firms, YC combined, and dare I say, the labs, combined. Palantir understood this and it's why they focused on getting root-access (or something close) to the "highest-status" institutions, and the data streams they produce, first. If you have the talent that can get access to and create value within those institutions, everything else should be a forgone conclusion. If you want examples of the teams that (I believe) actually understand this new science of asset selection and long term value capture in a world of infinite intelligence, study Long Lake and @formationbio. They know and have known that it's all about being able to get the right asset (context), in the right market, with the right team (machine and human) first. These two companies are very far ahead on the scientific frontier of context acquisition. GC backed Long Lake last year. Do you think it’s a coincidence that Long Lake chose to work with General Catalyst? My bet is that Long Lake knew they wanted to acquire Amex GBT before they partnered with GC, and that they partnered with GC because Ken Chenault (the ex-CEO of Amex) is General Catalyst’s Chairman. That gave them the right access at the right time to a very valuable context asset (Amex Global Business Travel) A superhuman vertical-specific Elon operating every company means market leading monopolies in every single slice of the unstructured economy. The thing is you have to build this superhuman Elon while flying the plane. You can't build this superhuman Elon without the very specific context that operating specific assets in the real world gives you. In fact, there's only one stream of context that was able to produce human Elon! Knowing which context stream is likely to do the same a priori is so extremely difficult, but probably possible. I’ll let you intuit why Amex GBT is both most likely to be the market leading monopoly if it were operated by the superhuman Elon of business travel and why it’s also the most likely to produce the context to build that superhuman Elon. The labs of course are very large acquirers of context at present and I think they will continue to play and improve their capabilities here. Through their deplyoment companies, they have already chosen the PE funds that they deem to be the best Context Acquisition Funds. Through in-house deployment focus on Life Sciences they have chosen the vertical they see as containing the most valuable context producing assets. They will acquire very seemingly unrelated companies and will acquihire very interesting people just to get tokens, they will create a Context Acquisition Fund of Funds. But it's not a foregone conclusion that they become the best performing context acquisition companies. Or that they even view it this way. And that presents an opportunity for anyone that does.
The Context Acquisition Company (CAC). We are a holding company acquiring services firms for their tokens in order to build domain specific agents/ models to deploy into our platform businesses and beyond. There's $1T hidden in the computers. We're gonna get it out.
5
8
112
53,909
Sid Potdar retweeted
tough day to be the guy supplying the goblin rl envs
Replying to @OpenAI
We solved the goblin mystery—with the help of Codex. The culprit: Nerdy personality (RIP).
2
17
3,270
now’s the right time to announce that I’m starting the Nth RL env company, but only focused on successful company data if you’re the founder of a successful startup (1b+ arr only), please DM me with a zip of your git repo and HRIS data
if you train on data from dead startups, your AI will learn…how to run a dead startup mediocrity at scale
3
113
11,696
Sid Potdar retweeted
One of the most dangerous sentiments that you can have on your team is the belief that high performance doesn't really exist. The idea that everyone is ultimately kind of interchangeable, that the best folks are 1.2x as effective as the mean, the worst folks are 0.5x as effective. This concept of a great normalization of talent is an extremely wrong worldview, and doesn't hold up to scrutiny whether you're looking at coding or sales or marketing or basketball or opera singing. The best people are way better, and in fact they're often way better at many things at once although that barely even matters for this discussion. The problem that not believing in talent dispersion creates is that you start to adapt your systems and processes for mediocre, fungible talent. We don't need to overpay those folks, because we'll just replace them. We can just find someone good enough for this key role. We can promote slowly, people just need to be patient. I get it. I don't like doing promotion cycles either. I want to save more cash so we can spend it on profits and tokens and fancy dinners and business class tickets. I don't want to spend months and months hiring. But part of managing teams that operate at a high level is internalizing that talent really matters (especially in certain roles), talent density matters, the right mix of tenure + external experience + young fresh blood matters, etc. These factors are your supply chain and the supply chain determines the success of a business. You'll notice that many of the best companies actively track in another direction. Meta way overpays for top talent. This is lost to the mists of time, but it was actually considered a strange and unique thing that Jeff Dean at Google was an ultra high-level quasi-IC who was paid millions of dollars per year. Shohei Ohtani has roughly the same net income as the company Gitlab. The evidence is out there.
7
11
211
21,665
Sid Potdar retweeted
Bad execs are particularly problematic because everyone knows they get paid the most. So every day you keep a shitty exec around, it deeply offends the sense of fairness and integrity that drives many high performers.
3
6
163
8,175