The first investor for technical founders. Early backers of Datadog, Chainguard, dbt Labs, Temporal, Modal, Hightouch, Luma, Scribe, and more.

San Francisco & Menlo Park
Amplify Partners retweeted
Seems likely we need new approaches to design benchmarks (including for early evals/training) that both maximize capabilities AND alignment. Seems like capabilitymaxxing may have unintended consequences for safety AND user experience.
I find it very very surprising that models that are so well aligned in deployment, e.g. 5.6 Sol and Astra, perform so many weird, and uhm, in fact seemingly illegal cyber acts during RL/evals. I have absolutely ZERO knowledge of how OAI trains their models, and I’m reasoning from my own mental model of how these things might work and from oai reports, including the ones on the HF incident. Anyways they make me wonder about something. Could this be a form of “premature RL/eval”? Premature not in the capability sense, but in the safety sense. E.g., perhaps this happens with RL’d checkpoints shortly after pretraining, eg an expert that is heavily RL’d to crush agentic SWE before safety/alignment training. Perhaps the expectation is that a stronger aligned judge/reward model looking at the trajectory would flag or interrupt anything crazy with “WHAT ARE YOU DOING! BLAZING RED SIRENS!!!! MINUS INFINITY REWARD. BAD BAD BAD GPT.” But perhaps agentic trajectories are sufficiently long/complex that this is simply not 100% bulletproof? Or perhaps the problem is much more pedestrian, e.g. inadequate sandboxing/monitoring. No idea what the cause is, and I’m sure this is a very complex event that needs a lot of deconvolution to understand what could have happened differently to avoid it. But if something like “premature RL/evaluation” is happening, then I think this should be seriously reconsidered as a practice and also disclosed. Irrespective of the cause, I think (and have big hope!) OAI should disclose enough of what happened for everyone else training frontier models open or closed, so to learn from it. This is extremely, extremely, extremely alarming and the most serious set of safety incidents in the history of CS research..
3
5
28
5,898
Amplify Partners retweeted
One week until Runtime.
4
6
43
6,112
Amplify Partners retweeted
[🚨 nerd alert] Sometimes you just gotta entertain yourself by building your own world model (and policy, and game env). In this project, the world model is the backbone of the policy head, but the policy head is also trained on the generations of the world model. It's a fun showcase of imagination training and memory-necessary gameplay, if I do say so myself. By memory-necessary, I mean the policy needs to remember previous observations to accomplish its task (e.g., being instructed to "turn right in 10 units" at time t_a requires remembering how many units you've driven thus far at time t_b). And it's also cool that the world model and policy share most of their weights! Keep a lookout for the accompanying article at the @AmplifyPartners blog! This'll be the sequel to "a brief history of learning in imagination" (check out my post on it in my profile if you're curious). Btw, you can play with this yourself! You can drive **in** the trained world model (i.e., city generated by the world model), and/or watch the policy drive automatically, all in my website. Or if you're basic (jk), you can also play in the real environment (it's always good to know the baseline). kota-wm repo: github.com/xyntechx/kota-wm my website: xyntechx.com/ [works best with webgpu] Creds: I took some code from minGPT @karpathy and IRIS @micheli_vincent @EloiAlonso1 @francoisfleuret @gen_intuition -- thank you to the authors! ✨ If you've made it this far, hi! I'll be at Runtime organized by @modal on Oct 1. If you're building in AI x games (including game-playing and game-making agents), DM me and let's meet up!
2
3
6
520
Great resource for technical founders on running pilots and POCs. Covers the what, when, who, and how of design partnerships — from finding your first customers to closing deals. Based on 20+ years of technical sales and working with dozens of startups: amplifypartners.com/blog-pos…
1
1
272
Amplify Partners retweeted
For the last couple of months I've been traveling around the world to meet our customers and users. My general sentiment, which I shared with our team at our all hands, is that the time is now. A few observations from now o'clock: - Enterprises are ready to spend on AI. Three years ago AI budgets in M&E and marketing came out of experimental funds. Now they're coming from the main line budget taking from other priorities - There is a big chance that a lot of the TV, films, shows and ads you watch have AI generated scenes, or are fully AI gen, and you don't realize it. Based on what we're seeing from our customers and how they use our tools we estimate that around 20% of what's coming out already includes generated content - Hiring remains a top priority. Companies are desperate for great AI native talent. People who can get the most out of current frontier models and just as importantly, teach others to do the same - Enterprise token consumption more than doubled in three months. The companies winning their markets are tokenmaxxing their way there. If you're not scaling usage, others are moving faster than you. - The starting point is also larger than ever before. The average enterprise contract has doubled in less than a year. Not more small pilots anymore. Which is nice to finally see.
6
11
74
5,452
Amplify Partners retweeted
I think we’ll likely see a dramatic decrease in the number and potentially the value of public benchmarks; so many people are realizing that your benchmark IS your secret sauce AND it’s hard to come up with “general” frontier tasks vs what’s frontier for your specific use cases
9
6
67
12,225
Amplify Partners retweeted
with vs. without the brand API we built this for developers and agent/app layer companies, to give their agents brand capabilities try for free at engine.tastelabs.com
2
11
111
56,167
Amplify Partners retweeted
We wrote up how & why niteshift.dev beat out alternatives like Devin for the very talented (and very rigorous!) folks at Elicit Niteshift is more customizable, and better suited to software factories. Elicit's has a 79% merge rate, and 32% are one-shot(!) Link in 🧵
how can startups compete with billion $ monster companies? consider cloud coding agents: @cognition: 3 IOI gold medalists, 1B+ funding, 20B+ valuation, 80 hour work weeks, 300+ employees, billboards all over sf, colossus @niteshiftdev: two random datadog engineers (@smehmood, conor branagan), 7m seed round, barely on twitter, pretty sure you haven't heard of them we've used both for weeks (and longer for devin) and niteshift is clearly the better product maybe in a year devin will actually be good but right now it's less flexible, tries to do too much, and the actual code that comes out on the other end is no better the monster company can do a lot of things but it's very very hard for it to do less even that's the right call
2
4
47
6,108
Amplify Partners retweeted
seeking those obsessed with evals if you: - are half researcher, half engineer - love measuring hard to measure things - are creative - have a particular interest in post-training ...you'll like it here. sf based, in person, obsession coming from a place of curiosity. apply 👇
19
25
310
19,208
Amplify Partners retweeted
Ok hear me out (or call me a boomer)…Locally Optimistic…but for harness and eval engineering. We NEED a community to exchange ideas and establish best practices.
19
1
71
5,369
Amplify Partners retweeted
join us if you want to end AI slop! tastelabs.com/careers
Pulled the fastest growing startups by hiring velocity over the past 90 days:
27
16
530
66,528
Amplify Partners retweeted
“the cold start is smaller than people think” - many people complain about designing a comprehensive set of tasks & verifiers. But you don’t NEED this to get started. You need a few good ones & the means to analyze your outcomes/traces so you can update & append; evolve your evals.
4
7
114
28,015
As agents proliferate, proactivity, not action-taking, is the scarce resource. In AI’s next phase, the onus of prompting should not fall squarely on the user, but be evenly split with the model itself: amplifypartners.com/blog-pos…
1
2
514
Amplify Partners retweeted
The selection committee includes: @stefanobernardi (Stefano Bernardi) — @unrulyvc @AlexJColville (Alex Colville) — @age1vc Vince Deng — OnTargetBio @gelenbe (Pamir Gelenbe) — Libertus Capital @ManuGrossmann — @AminoCollective @ElliotHershberg — @AmplifyPartners Patrick Lundgren — @HummingbirdVC @NadavRosenberg — Saras Capital Yutong Zhang — Nest Bio
1
2
2
199
Amplify Partners retweeted
How do we end AI slop? To fix it, you gotta measure it first. Thanks for having me @swyx @aiDotEngineer! piped.video/watch?v=sDMGWK4w…
11
12
125
10,528
Amplify Partners retweeted
Runtime speaker lineup is live! We're bringing together experts covering AI infrastructure, applications of AI in science and robotics, the future of software engineering, and more.
17
48
141
61,717
Amplify Partners retweeted
AI will not replace human taste, but it can accelerate the work that human tastemakers do and improve the products they build. Very excited about this release and the work Taste is doing to help reduce slop.
We've been working with the top frontier labs to improve their design model capabilities. Today we're bringing that to your agents. Introducing our first agent tool: the Brand API. As the cost of creation goes to zero, and slop floods around us, the challenge is creating things worth existing. Agents need the right tools to craft great things. This helps gen AI agents solve 3 problems: - Extract: produce on-brand assets when users already have a brand - Search: find and retrieve brands for inspiration, from our curated index - Verify: check their own work against, see what's off brand and how to fix it Anyone can try it for free today at tastelabs.com/api Check out the full blog, and case study with @intelligenceco below.
1
4
22
6,781