This is exactly why Spiral is focusing on AI data infrastructure. Training on the right data is a surprisingly outsized lever
Pretraining progress seems to be coming mostly from data improvements. @who_is_jerbear and I pretrained combinations of year-representative open model recipes and data corpuses across 2019 to 2025 at various small scales. Data improvements contributed 3.24x as many compute multipliers as model improvements did (12.0x vs 3.7x). And the gains stack independently - a better dataset helps every architecture about equally, and vice versa. Here are full results, plus what we think this means for the future of AI progress: dwarkesh.com/p/pretraining-p…
1
7
604
You heard it here first folks - @vortexdotdev is the solution to the world's hardest problems
Replying to @OpenAI
The group produced an analytical proof and Lean formalization that via Navier-Stokes dynamics a fluid can develop a singularity in finite time. The solution is a vortex, a spinning swirl of fluid, that spirals inward and gets increasingly elongated, like spaghetti. cdn.openai.com/pdf/32d9f210-…
6
401
I thought I'd test out my theory by vibe coding a Vortex plugin to read Parquet files... It works!
I’ve written enough file formats to know the bytes are the easy part. I/O, pruning, scheduling, pushdown eat the weekends. I built Vortex so I never have to do this again! spiraldb.com/blog/vortex-the…
1
2
14
2,518
I’ve written enough file formats to know the bytes are the easy part. I/O, pruning, scheduling, pushdown eat the weekends. I built Vortex so I never have to do this again! spiraldb.com/blog/vortex-the…
4
18
152
17,256
Four innocent lines of PyTorch hide a video preprocessing footgun. In our H100 benchmark, the resulting physical plan made the model’s first operator run 8.9× slower. The code is clean. The physical plan isn’t. 🧵
1
2
13
2,261
Fusion doesn’t always win though. At 64 output channels, the matched benefit falls to 1.2×: the fused schedule repeats preprocessing, while materialized RGB provides reuse.
1
3
80
Write the pipeline you want. Let Spiral choose the physical plan. Materialize RGB, keep native YUV, or fuse preprocessing into the model. Let the optimizer decide. See how @SpiralDB avoids the 9× footgun: spiraldb.com/blog/when-video…
1
4
83
Spiral's optimizer is getting quite good these days
9
1,411
Ohhhhh, they were running the frontend off an OLAP database. Makes sense now.
We launched Rippling Data Cloud today - an all-in-one rebuild of the modern data stack, with AI deeply integrated throughout. Why would you want an org-and-employee-centric data stack? Well, here’s how I used Rippling Data Cloud to help with token burn and cut AI slop. 1/ nitter.net/parkerconrad/status/20…
4
632
SmithDB is the perfect example of how far performance can be pushed by having full control over the storage layer. The DataFusion + @vortexdotdev stack seems to be emerging as THE way to build next generation databases. #ParquetIsForFloors
We built SmithDB: the database purpose built for agent observability workloads that now powers many parts of LangSmith. Agent observability presents a challenging data problem. Agent traces can contain tens of thousands of intermediate spans and large, unbounded payloads. These characteristics are a direct result of agents running for longer time horizons and LLM context window sizes growing. Traditional data infrastructure was not built to handle the complexities associated with storing and querying this data. SmithDB brings LangSmith up to 12x performance improvements across access patterns most important for agent observability. I’ve been working on SmithDB directly with an amazing team over the past few months, and I’m incredibly proud of the results we’re seeing. I wrote a bit more about the story and engineering challenges behind SmithDB in this blog. Additionally, if you’re a systems engineer interested in building the future of agent observability, please reach out!
2
2
47
5,598
Nicholas Gates retweeted
Love seeing Naomi Osaka honor the CLRS Algorithms textbook at this year's Met Gala
108
1,700
17,879
621,592
Vortex is BOTH 38% smaller and has 10–25x faster scans than Parquet+ZStd for TPC-H SF10 We implemented BtrBlocks-style cascading compression in Vortex that recursively tries codecs like ALP, FSST, and bit-packing, letting the data pick the best encoding spiraldb.com/post/cascading-…
7
41
2,913
🚀
Checkout our latest post by @andreweduffy about bridging our Rust code to Java to accelerate Apache Iceberg queries 🏎️ spiraldb.com/post/vortex-on-…
1
495
Slightly late to the party here, but could the billion dollar #ErasTour really not figure out how to center a digital clock…
5
426
We can't be the only ones who compete with `cargo clean`
5
321
I bet there aren't any other file formats that come with a slick terminal explorer 😎
1
21
1,485