Pinned Tweet
I find the Lakebase design for serverless Postgres very elegant, so I spent some time explaining how it works in this blog. The blog starts by explaining how databases really persist data (with a write-ahead-log and data files that are updated async), and how Lakebase separates storage and compute by externalizing those two components. It ends with how the Lakebase architecture naturally leads to LTAP, enabling OLTP and analytical workloads against a single governed copy of data. My goal was to make it readable by anyone curious about how these systems work, not just database and storage experts. That turned out to be a lot more challenging than I first thought. Database storage is one of the most complex areas in computer science (the ARIES paper cited in blog was the hardest paper I personally ever had to read). The first draft had too little detail and I couldn't land the ideas. The second had too much and I'd lost anyone who isn't already a storage expert. This is the third draft, and I'd love feedback on whether the depth feels right. databricks.com/blog/lakebase…
7
29
164
31,788
Reynold Xin retweeted
We’re bringing Databricks Unity Gateway to developers on Neon (@neondatabase), and it’s the fastest AI Gateway for Kimi K3! Databricks inference is moving super fast!
6
11
114
11,284
Reynold Xin retweeted
The results for AI Gateway Latency benchmarks are in. Here is one for @kimi_moonshot. 🥇@neondatabase 🥈@blazerail_ai 🥉@ngrokhq
1
4
13
3,156
Reynold Xin retweeted
The Neon backend is GA. A complete set of primitives for apps and agents built around the database, with open standards, no lock-in, and branching-first by design. Starting with: Lakebase Postgres, Object Storage, Functions, Managed Better Auth, and AI Gateway. More to come.
24
22
136
27,493
Reynold Xin retweeted
Built a full-text search over all of Hacker News (~49M stories, comments, jobs & polls, syncing live per hour) to be 𝐬𝐮𝐫𝐩𝐫𝐢𝐬𝐢𝐧𝐠𝐥𝐲 𝐟𝐚𝐬𝐭 with no Next.js level cache and all magic at the Neon Postgres layer: hnfaster.vercel.app 🎉 • Every keystroke is a live query against @neondatabase Postgres • BM25 (powered by Lakebase Search) ranking runs inside the Postgres • No connection pool, just SQL-over-HTTP round trips • pg_prewarm loads the table + every index into cache • partial BM25 index per type, predicate matching the WHERE clause • Deployed on @vercel Just Postgres doing Postgres things, close to the compute 🐘
18
11
202
55,611
The best database for agents and recommended by agents is Lakebase Postgres on neon.com/
#1 in Claude Code, Codex and Cursor. Still first for GPT-6 Astra and Fable 5.1 @neondatabase ripped our leaderboard of which database agents choose. Agents cite easy setup, branching and lower compute bills through scale-to-zero. Every run open 👇 armature.tech/leaderboards#a…
4
13
83
13,203
Pretty good summary!
Sunday morning reading - a @VLDBconf 2026 paper on Lakebase from Databricks .. a very nice overview of the architecture that belongs to the same family as any system that treats object storage as the durable backplane and puts a low-latency cache/log in front of it. Here are some points on how Lakebase handles durability and failure handling, things that differ so subtly across similar platforms like Aurora .. • Commit point. Lakebase does not fsync WAL on the Postgres compute node. COMMIT returns when a quorum of safekeepers - typically two of three, in different AZs - has flushed that WAL record to local disk. Pageservers and object storage are not in that handshake. • What each layer is allowed to lose. Compute can vanish at any time, it holds no durable state. A pageserver is a cache, so losing its NVMe does not drop committed transactions. Durability of recent writes sits on safekeeper disks until those segments are uploaded to object storage. • The tail window. After SUCCESS, new WAL exists only on the safekeeper quorum. S3 catches up later, tracked by backup LSN. If a quorum of those flushed SK copies is destroyed before that upload, acknowledged writes are gone. That is the window, it is not closed at commit time. • Compute and pageserver failure. Compute crash, scale-to-zero, and failover are 0 data loss because a new compute attaches to the same log. Pageserver NVMe loss does not drop committed transactions, another pageserver can reconstruct pages from WAL and from layers already in object storage. • Object storage as the long-term copy, not the commit copy. Once backup LSN has passed a segment, Safekeeper-local WAL can be dropped and history lives in the lake. Branching and Point in time recovery use that LSN-addressable history. They do not move the commit rule: SUCCESS still means quorum disk flush, not an S3 PUT.
1
35
5,881
Lakebaseの論文がVLDBで公開されている!AuroraやAlloyDBは第2世代で、オブジェクトストレージ上にオープンな形式でデータを保存するLakebaseは第3世代ということらしいよ。 vldb.org/pvldb/vol19/p4385-p…
2
17
100
12,789
Reynold Xin retweeted
#VLDB2026 is underway, and Databricks is here as a Gold Sponsor. We’re sharing what we’re building for the future of agent-native data infrastructure, including: • a keynote from Databricks co-founder @rxin on why the agentic future demands Lakebase Postgres • AutoLiquid and Ultron query optimization • Lakehouse//RT, Lakebase, and LTAP • conversations with the database community on the ground Still ahead: sessions and demos on Lakebase, @ApacheSpark Structured Streaming, and Enzyme. If you’re at VLDB, come find the team 👋
3
8
49
5,840
Reynold Xin retweeted
Postgres on S3 is the right design, but even more interesting is making WAL the source of truth. This is how our storage works - we wrote a deep dive, summary on thread: neon.com/blog/wal-s3-lakebas…
12
27
326
114,114
Reynold Xin retweeted
This is true. It wasn't actually possible before 2010 because datacenter networks would bottleneck. We used to design coupled storage/compute where you brought compute close to "big data". But research on "full bisection bandwidth" networks made it possible to essentially just talk from any machine to the storage system at full speed. The disaggregation started then! Databricks and Snowflake started soon after many others followed. Now "Put it on the object store" is the way to go.
"Put it on the object store" remains undefeated. - OLAP: Lakehouse - OLTP: Lakebase - Kafka: Warpstream - Vector Search: Turbopuffer - Git: Cursor Origin
9
27
341
46,480
Reynold Xin retweeted
Today, we announced that we crossed $7B in revenue run-rate, growing over 80% year over year in Q2. We also shared: 🚀 $100M+ revenue run-rate for Lakebase 🚀 $1.5B+ revenue run-rate for Lakehouse, growing over 100% year over year 🚀 Continued positive adjusted free cash flow And we raised $5B in our latest fundraise. We’ll use this capital to invest in: 1️⃣ Lakebase, our serverless Postgres database built for AI agents 2️⃣ Genie, our AI coworkers that actually understand your business data 3️⃣ Unity AI Gateway, our multi-AI governance solution that helps control costs @iamVictorDey shares more in @Forbes: forbes.com/sites/victordey/2…
76
243
1,012
308,384
Reynold Xin retweeted
Excited to share that we've acquired ElectricSQL, the team behind PGlite. Agents need super fast Postgres and this team built an amazing WASM (WebAssembly) implementation of postgres that runs in your browser, but can sync back with Postgres instances asynchronously. Exactly what blazing fast AI agents today need. Excited to supercharge our 𝐋𝐚𝐤𝐞𝐛𝐚𝐬𝐞 𝐏𝐨𝐬𝐭𝐠𝐫𝐞𝐬 offering with these capabilities. databricks.com/blog/electric…
20
72
461
67,731
Reynold Xin retweeted
Replying to @databricks
@databricks has the fastest and lowest latency on Kimi K3 (max)! Great job by Databricks AI team! artificialanalysis.ai/models…
17
249
749
26,785
Reynold Xin retweeted
Here’s another awesome use-case for what we can now do with our unstructured data because of AI agents. Box now works with Databricks so you can take structured data from enterprise content (like contracts, financial document, supply chain data) and connect that data into Databricks. This means that I can now query large document datasets without moving or reprocessing that content. And you can connect the data with any other system, like your ERP data, CRM, or product analytics. This opens up a ton of new use-cases for enterprise content. All possible because of headless software and agents.
The Box MCP server is now available in the @databricks Marketplace. Data analysts and data scientists can now pull unstructured content from Box directly into their Databricks workflows through Genie. Structured vendor metrics and contract terms in one place. No platform switching. No data duplication.
20
23
159
51,743
Reynold Xin retweeted
Gartner’s Magic Quadrant for Analytics and Business Intelligence (BI) is out, and Databricks was named a Visionary in our first appearance, the highest debut for any vendor in this MQ’s 20+ year history. BI has already changed. Anyone can drill from a signal down to the truth behind it and agentic loops make sure every answer has the full story. That’s where we're ahead with Genie and AI/BI. I use it every day to see the key signals, understand what changed and why, and decide what to do about it. Huge congratulations to the teams, and thank you to our customers.
8
34
147
15,951
Reynold Xin retweeted
We benchmarked coding agents on our own internal tasks at Databricks and learned a lot! There are many surprising opportunities to lower cost and increase quality, and many models including open source ones are truly competitive now. 🧵
70
153
1,067
256,527
Reynold Xin retweeted
Happy 30th birthday Postgres! 🎉 What an accomplishment, almost nothing in tech has stayed relevant for 30 years. From humble beginnings as a Berkeley research project to becoming the default for everything from side projects to massive production systems. 30 years later, the database and the community have more momentum than ever. It's why we built Neon on Postgres, serverless, but 100% the Postgres you love. 🐘
2
14
78
34,816
Reynold Xin retweeted
My co-founder @rxin personally wrote this really good paper that explains the main idea behind postgres Lakebase as well as LTAP. It almost serves as a primer on how transactional databases are built and how Lakebase and LTAP work. Maybe more importantly, what are the tradeoffs, and what are you giving up by adopting this new approach. Highly recommended reading: databricks.com/blog/lakebase…
9
72
336
31,971
Reynold Xin retweeted
You may have heard that GLM-5.2 at 328 token/s is cool, How about 392? Databricks is now #1 in inference speed for GLM-5.2 on Artificial Analysis. It's a great model, and we did a lot of optimizations.
93
93
1,199
324,494