Engineering Lead for Agentic Infrastructure @databricks | Scout @greylockvc (angel investing in data/ML/cloud) | prev: @stanford @ucbrise

San Francisco, CA
Ankit Mathur retweeted
I now have more personal agents than flights to book
5
24
284
10,717
Lots of discourse online about Lunamaxxing and Opus 5.5, but we wanted to back our intuition in evidence. We used @databricks Unity Gateway to collect data based on our internal developers, and it shows the pareto frontier has moved! Many more exciting models coming this week!
Crazy few weeks for model releases! Our findings @Databricks show several new models meaningfully advance the pareto frontier. Results below (online workload analysis of N=2,400 engineers, plus offline evals): 1. Two of the three models released last week clearly expand the cost/quality frontier: Opus 5.5 and GPT-6 Luna. 2. Opus 5.5 is now the highest quality mid-tier model. It is better than all prior Opus models, better than GPT-6 Sol, and better than GPT-5.6 Sol. 3. Opus 5.5 reduces same-task costs consistently by 20% in both offline and online analysis. This is against a baseline of Opus 4.8, the prior least-cost Opus model (Opus 5.0 was a bit of a dud with high costs and barely noticeable quality improvements). 4. Due to best-in-class quality and lower costs, Opus 5.5 is a strong candidate as an “every day default” model for coding, and we are now encouraging it for this purpose at Databricks. 5. GPT-6 Luna is very, very, very cheap. It was at least 20 times cheaper per-task than Opus 5.5 in every offline benchmark we tested and in observed online use. 6. GPT-6 Luna is surprisingly capable given how cheap it is. On one of our most difficult evaluation suites it roughly matches Opus 4.6 performance, while being 99.3% cheaper per-task than Opus 4.6 was at that time. That's a 100X cost reduction in ~9 months! This finding is preliminary and we are still evaluating Luna quality on a broader set of offline and online tests. Our production setup: Unity Gateway to route workloads across models and trace agentic interactions. A mix of end-user harnesses including: Omingent (meta-harness), Claude Code, Codex, and Cursor.
1
2
18
1,336
We use this the CLI every day at @databricks to use every coding tool. That's how we end up using frontier models as soon they come out!
A new model comes out roughly every five days. The best model for a coding task today may not be the best or most cost-effective choice a few months from now. The new Unity Gateway CLI helps organizations keep up with that pace. Developers can keep working in Claude Code, Codex, OpenCode and other coding agents, while admins centrally manage models, tools, policies and budgets in Unity Gateway. When a better option emerges, teams update the configuration once and the CLI rolls it out across supported agents. Now with the Unity Gateway CLI, teams can: • Set default models, MCP servers, skills, Smart Routing and spending policies centrally • Launch approved coding agents with one command and automatically apply organization settings • Track spend and capture traces across supported coding agents databricks.com/blog/deploy-a…
5
24
1,694
Ankit Mathur retweeted
We rolled out OSS models to all our internal coding agents through Databricks AI Gateway. Then engineers *surprisingly* discovered how powerful they are. “I didn’t reach for Sol (or Opus/Astra) at all today.”
24
9
168
10,431
Ankit Mathur retweeted
All seven Databricks co-founders began our journey together at UC Berkeley. We started the company in a room in Soda Hall, moved to a small office on Addison Street, and spent our first few years in Berkeley before eventually heading to SF. Two of our co-founders are still on the faculty. Today we announced Databricks Field at California Memorial Stadium, our first collegiate athletics sponsorship. Berkeley shaped everything about how this company thinks, and this is our way of investing in the next generation of students and builders who'll do their best work on that campus. Go Bears! databricks.com/company/newsr…
126
222
2,342
225,392
Ankit Mathur retweeted
for those monitoring the situation this seems to be the official US department of education account doing a dogwhistle about getting rid of some Indian kids attending a UT football game
MAKE COLLEGE FOOTBALL GREAT AGAIN.
422
905
15,422
993,118
Ankit Mathur retweeted
Long any speech and debater
Dwarkesh at 17 years old btw
10
7
352
101,280
Databricks inference team is cooking!
The results for AI Gateway Latency benchmarks are in. Here is one for @kimi_moonshot. 🥇@neondatabase 🥈@blazerail_ai 🥉@ngrokhq
3
29
1,575
Ankit Mathur retweeted
The results for AI Gateway Latency benchmarks are in. Here is one for @kimi_moonshot. 🥇@neondatabase 🥈@blazerail_ai 🥉@ngrokhq
1
4
13
3,187
Ankit Mathur retweeted
At Databricks, more and more engineers are shifting to OSS models as daily drivers, barely reaching for Claude or Astra. Our inference customers are seeing the same thing: shifting even 20% of internal coding traffic from Claude/GPT to OSS can dramatically cut spend. OSS adoption in the enterprise is going to be massive.
60
69
1,039
108,428
Today @neondatabase Auth, Object Storage, Functions, AI Gateway are GA. Connected to Lakebase Postgres. Give it to your agents 𝚗𝚙𝚡 𝚗𝚎𝚘𝚗 𝚒𝚗𝚒𝚝
The Neon backend is GA. A complete set of primitives for apps and agents built around the database, with open standards, no lock-in, and branching-first by design. Starting with: Lakebase Postgres, Object Storage, Functions, Managed Better Auth, and AI Gateway. More to come.
3
10
57
5,157
Ankit Mathur retweeted
Ali is a big reason why I joined Databricks. I was planning to build another startup. Then @alighodsi kept calling me. He’s honestly the best founder I’ve met, and I’ve met a lot. I was like, “Shit, if I want to build another startup, I should learn from this guy.” I haven’t regretted joining Databricks. I'm still learning a lot from Ali every day.
Ali Ghodsi never wanted to be CEO. In 2015, he was interviewing for a professor job at Berkeley when the @Databricks board handed him the interim title. Revenue that year was $1.5M. This episode with @alighodsi will go down as one of my favorites. 10 things I took away: 1. Focus the entire company, orders of magnitude of attention, on its single biggest bottleneck. Like a laser, almost to an extreme. The cycle is 1-3 years, not weeks. If your focus changes weekly, then you’re just in firefighting mode. 2. There is nothing worse than a conflict-averse CEO. They are wonderful people, but they are in the wrong job. Conflict is the gym for a CEO: nobody likes it, but everyone has to go. 3. Study your enemy carefully, understand their weaknesses and apply your strengths to those weaknesses. Snowflake had 2x his revenue. He didn’t copy them. He found three weaknesses (proprietary, no AI, expensive) and hammered them account by account for four years. Watch the competition, never follow it. 4. The concept of a Lakehouse was ridiculed internally and online. No one wanted to market with this new term. So he made the whole company religious about it anyway, killed the ads that converted better without the word, and put it in the sales comp plan. It worked. All hands on deck, no exceptions. 5. Be willing to take a step back for a much bigger vision, even when the company is already succeeding. At multiple hundreds of millions in ARR, he was unhappy, because the vision he pitched investors wasn’t the company he was running. So he took one step back to go ten forward. 6. On the flat org, player-coach model that a lot of people have talked about this year: “it’s BS.” Separate how the company thinks (AI, ontology) from how the humans get managed (they are after all, still humans). His staff meets 3x a week. I asked if it could just be coordinated in a Google doc. His answer: do you meet your wife and kids, or coordinate that in a Google doc? 7. His test for a sales leader: can they build the car, or just drive it? Ron Gabrisko, the Databricks CRO, had seen $0→50M and $50→100M+, and hadn’t changed jobs in 10 years prior to joining. Now he has been the CRO for over a decade. He built and drove the car the whole way. That almost never happens. 8. Hire execs ahead of the curve because by the time you need them, it’s too late. A real search takes 6-12 months. The extra time helps you increase false negatives and decrease false positives. Do an insane number of backdoor references because 80% of ‘front door’ references are bs. 9. The best salespeople are not super technical, so stop trying to force them to be. Square peg, round hole. The best win with professional aggression, high EQ, and mapping the real power base (how decisions get made high up in an organization), not technical depth. 10. Yes, your best AEs will annoy people. One of the first at Databricks got a meeting nobody could get, but got banned from the customer’s building for it. He told Ali, “what are you complaining about? I got the meeting.” Professionally aggressive is the bar. One bottleneck, zero wussing out. He reminds me of @elonmusk that way. Chapters 0:00 – Introduction 1:22 – The secret CEO search and why the board bet on a founder 4:11 – Professor or CEO? Always taking the harder option 8:18 – Pour everything into one bottleneck 12:16 – Killing PLG and learning what great enterprise sellers actually have 19:09 – Hiring ahead of the curve: sales leaders, execs, and back-door references 27:02 – The Snowflake rivalry: study your enemy, never copy them 33:22 – Lakehouse: conviction, ridicule, and the case for second acts 41:02 – The killer instinct and why conflict-averse CEOs fail 44:10 – Dunbar's number and rethinking the org chart around AI 48:00 – AGI is already here — enterprises just use it as a chatbot 52:55 – Does he still code? Two days for a connector vs. three quarters 57:18 – A day in the life, and why the Monday meeting isn't theater 1:04:53 – Why Databricks will go public, just not yet 1:08:09 – Get over conflict aversion, or don't be CEO 1:11:07 – Brian's takeaways Link to more in the comments.
16
15
262
41,698
Ankit Mathur retweeted
Databricks rolls out Astra and learns that it outperforms the previous best models on long horizon tasks.
Today we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: 1. Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks. 2. Engineers given Astra increased overall coding spend by around 60% compared to baseline. 3. It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models. 4. We learned above by piloting Astra with around 200 users to gain signal on both quality and cost. We use Unity Gateway to do cohort-based experiments for all new models. 5. We give engineers a sub-budget specific to Astra to encourage them to use Astra selectively on complex tasks while preferring lower cost models for everyday tasks. Our engineers are able to mix-and-match tools and models within their overall budget envelope (we also allow for increased budgets through various mechanisms). These budgets are defined in Unity Gateway and regularly revisited. Note: We do not have robust comparisons of Astra-vs-Fable because we have net yet rolled out Fable widely due to data retention policies.
2
30
2,895
Ankit Mathur retweeted
wall-to-wall deployment of astra for engineers at databricks:
Today we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: 1. Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks. 2. Engineers given Astra increased overall coding spend by around 60% compared to baseline. 3. It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models. 4. We learned above by piloting Astra with around 200 users to gain signal on both quality and cost. We use Unity Gateway to do cohort-based experiments for all new models. 5. We give engineers a sub-budget specific to Astra to encourage them to use Astra selectively on complex tasks while preferring lower cost models for everyday tasks. Our engineers are able to mix-and-match tools and models within their overall budget envelope (we also allow for increased budgets through various mechanisms). These budgets are defined in Unity Gateway and regularly revisited. Note: We do not have robust comparisons of Astra-vs-Fable because we have net yet rolled out Fable widely due to data retention policies.
82
31
1,063
236,341
Every company is going through the journey of rolling out the frontier - check out @pwendell 's deep dive on what we saw with Astra
Today we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: 1. Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks. 2. Engineers given Astra increased overall coding spend by around 60% compared to baseline. 3. It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models. 4. We learned above by piloting Astra with around 200 users to gain signal on both quality and cost. We use Unity Gateway to do cohort-based experiments for all new models. 5. We give engineers a sub-budget specific to Astra to encourage them to use Astra selectively on complex tasks while preferring lower cost models for everyday tasks. Our engineers are able to mix-and-match tools and models within their overall budget envelope (we also allow for increased budgets through various mechanisms). These budgets are defined in Unity Gateway and regularly revisited. Note: We do not have robust comparisons of Astra-vs-Fable because we have net yet rolled out Fable widely due to data retention policies.
12
1,039
Ankit Mathur retweeted
The "whiteboard defense:" I should be able to pull you aside at any moment and ask you to explain any customer-facing system you've shipped. You should be able to clearly explain how it works and defend the decisions you made. This is my benchmark for responsible AI usage. I don't expect line-level familiarity with the code. I don't care if you remember the exact function name or implementation detail. You may not even know it. I don't care. But if I ask "why did you do X instead of Y?", "what happens if this actor behaves maliciously?", "what data structure did you use here and why?", or "where does this fail?" you should be able to answer confidently. For PoCs, demos, experiments, whatever: I don't care. Generate 100% of it and understand none of it. Speed over quality every time in those specific scenarios. But if you're shipping customer-facing work, you can't be shipping things you don't understand at a high level.
231
980
9,332
477,641
Ankit Mathur retweeted
It’s been 23 days since I last opened Claude Code. It used to be one of my favorite products, but I think I prefer Codex now. Codex just works better with OSS models. And among OSS models, Kimi K3 is still unbeatable at coding.
76
32
842
59,180
Ankit Mathur retweeted
This entire discussion is just so ludicrous a) if you believe it is existentially dangerous, put in actual controls with a real regulatory body. Lock that shit down. b) if you don’t, chill the fuck out. Or treat like the Internet c) there is no c I’m solidly in (b).
79
71
963
76,964
Ankit Mathur retweeted
5 months at Databricks. I do not say this lightly. Our AI inference is getting too fast. We are extremely close to the model answering before you finish typing the prompt. We are not asking for a ban. We are asking for a pause.
152
67
2,447
219,961
Ankit Mathur retweeted
Steve Kerr says Roger Federer taught his Warriors team the secret to sustained success wasn’t outworking everyone, it was building a life around your craft you never want to escape from “I’ll tell you a great story about Roger Federer. We were playing in China with the Warriors in 2017. Roger was in Shanghai for the Masters tournament, so we invited him to come speak to the team.” “Draymond Green asked him, ‘How do you sustain success? How have you managed to win majors 20 years after you first won one?’ Everyone was expecting him to say, ‘I work harder than everybody,’ the will, all the clichés.” “Roger Federer says, ‘I get up every morning and I make breakfast for my kids. Then I take them to school and drop them off. Then I go practice tennis for about two hours, and I’ve figured out a really good routine where I don’t destroy my body, but I can get all the work in I need.” “‘Then I go have lunch with my wife. In the evening, we cook dinner together. The kids are all home from school. We ask them about school. There’s so much joy in the house. We put them to bed, and then I put my head on the pillow and go, “Man, that was just a great day.” And I’ve been doing that for 20 years.’” “And it was like, yes. That’s it. That’s the formula. That’s what leads to sustained success. He loves tennis. He loves his family. He loves life. It’s not outworking everybody and banging your head against the wall, ‘I’m going to be better than everybody.’ It’s allowing your natural talent to shine through with a work ethic, with a great family life, with perspective, with peace, with mindfulness.” “It’s perspective. It’s joy. It’s passion. It’s mindfulness. It’s an awareness that we are truly lucky. To get the most out of ourselves, it’s our daily rituals and experiences and love and joy.”
225
3,277
23,969
2,550,892