Pinned Tweet
I’ve started a new company with @tkkong! TK is a driving force behind a lot of Ramp’s success, building much of the core product, incubating the procurement platform, and leading Ramp Labs. We’re a team of IMO and Physics Olympiad gold medalists, and we’re hiring the most talent-dense team.
I’ve started a new company with @philhchen! Phil built frontier LLMs across research & engineering at OpenAI, DeepMind, and Scale. I was shipping AI experiments at Ramp Labs. We've been heads down building personalized AI coworkers for every business. We’re growing our team of researchers, designers, and IMO gold medalists. Reach out if you're interested!
66
16
417
175,932
Phil Chen retweeted
Interesting work by Mecka. We’ve gotten to know Josh and @jasontheutopian this year and they’ve assembled a cracked team for robotics. We’ve also seen step function improvements with latest models like Astra/Opus-5.5 so benchmarks here are increasingly important!
Are robots ready to accelerate science labs? We created the WetLabs Benchmark to test 9 critical tasks, from easy to hard, and let 3 leading models take 20 attempts per task Astra leads but it comes close...
1
2
15
1,265
we're hiring. our team has a former millennium prize problem solver
20
6
686
61,244
Today shows that alignment is as much a human incentives problem as it is a research question. We haven’t solved human alignment yet either, but capitalism + democracy + diplomacy has been mostly working
3
2
24
2,725
👀who's distilling whom?
16
3
338
41,943
We built ConceptCool for ourselves and it turned out too good not to share. Multiplayer app-building with full @Railway underneath. You and a teammate edit the same prompt live, and the app you get scales from 1 to 1000 out of the box. It has real queues, containers, and background jobs, so it actually reflects real tools beyond a Claude artifact or CRUD app. Nothing we've used comes close for shipping internal tools fast. Go try it.
2
2
12
3,288
model routers today are solving the wrong problem.

the goal of a router is to select the model with the best ROI for solving your problem.

the issue with router companies today is they don't know how you measure your ROI. the same "solve Riemann hypothesis, you can do it!" prompt might matter 1000x more to one company (any frontier lab) than me.

for many products, latency actually impacts user satisfaction more than response quality. routers trained only on eval quality vs cost miss this as well.

in an ideal world, every company should have their own model router, but most companies probably won't be able to quantify their ROI.
16
4
69
24,699
Codex computer history is pretty cool. Each event is stored as a JSON with the AX (accessibility) tree representing the contents of the window. These events are grouped into events.jsonl and metadata.json files for every 10-minute window. Codex summarizes these events into 10-minute and 6-hour md summary files. Presumably the 6-hour md file is just a summary of up to 36 10-minute summary files. Overall neat design and think this could help users discover more agentic use cases.
We've gotten a lot of great questions on privacy implications of Computer History. Here's a few things we've done to build a super powerful feature while keeping things private. First off, you can review all of your Computer History in the timeline view:
4
17
3,808
The service seems to be called "Skysight", referencing sky.app/, @AriX's company that was acquired by OpenAI
1
3
626
I think all current long-horizon agent designs are wrong because they underestimate the jaggedness of agent intelligence. in a jagged-intelligence frontier, agent productivity is bottlenecked by the tasks models can't do, whether due to access or skill. long-horizon enterprise agents need a richer interface for interacting with human colleagues. today agents approximate that by asking questions, but that doesn't solve the skill or access issue. for example, if your task involves uploading an SD card of data and training a model, the agent should be able to hand that off to a human. if the agent's training run attempt goes awry, the agent should be able to ask for help.
3
23
2,159
Phil Chen retweeted
Foundry is one of few companies well positioned to solve the physical AI data problem with their flywheel. They actually deploy robots, collect data, improve the robots, and then manufacture more. I'm grateful to support some of their operational workflows with Concept.
1
2
21
3,923
Phil Chen retweeted
We’re hiring for designers & engineers. Come work with me & @philhchen
88
7
443
32,221
anyone else notice this weird TTC scaling for Opus 5 on FrontierCode?
4
21
2,496
We’re hiring a designer. If you or any of your designer friends want to work on bridging the gap between frontier capabilities and intuitive UX, please reach out! @tkkong is a design legend and I’ve learned so much from him
2
4
24
5,684
We've built something our users love - DM @tkkong or me to try it out
We're opening up early access. Phil and I have been working with teams to ship internal AI apps and agents. We’re applying our learnings from Ramp Labs, OpenAI, and DeepMind to help operations teams scale faster. DM if you're interested in trying it out!
4
1
21
5,742
> saw some new code in our codebase that explicitly only supports Claude models > thought to myself how much I hate Claude models for writing code like that > mfw I learn that 5.6 Sol wrote it
6
43
6,100
Nothing like using 5.6 Sol to use web browsing to send a customer support ticket to AWS, receiving some AI slop back from AWS support, and then using 5.6 Sol to review the slop and reiterate the need for a human (or smarter support model) to review the ticket
8
2
38
4,750
Sonnet 4.5 is around 100 layers? transformer-circuits.pub/202…
19
3,365