Distributed query engine providing simple and reliable data processing for any modality and scale (github.com/Eventual-Inc/Daft)

A teleoperated robot episode doesn't start when the recording does - the operator is still setting up while the camera rolls. The robot's own joint positions tell you which frames matter without decoding a single pixel. With this method, trimming all of DROID - 500 hours of robot data - took 32 seconds on a laptop. daft.ai/blog/cutting-dead-fr…
6
332
Hand-pose annotation is the plumbing between raw robot video and a trained model. daft-physical-ai packages it into a single Daft UDF. This is how it works: track_hands in daft-physical-ai passes a Daft image column and gets a hand-pose column back. MediaPipe (CPU/2D) or WiLoR (GPU/3D MANO), same schema either way. Lazy, batched, distributed — Daft handles execution. On EgoDex: detect=100%, PCK@.1/.2/.3 = 49/84/96. Try it: pip install daft-physical-ai Check more use cases here eventual.ai/blog/announcing-…
1
2
300
Excited to announce: daft-physical-ai, a new Python data processing library for physical AI. We're starting with two use cases: hand tracking and reward scoring - essential steps for making robot data useful for training. daft.ai/blog/announcing-daft…
4
518
Some founders find their calling at 30. Sammy found his at 12.. and it led to Eventual. Why is he the one to do this? Number one: he loves robots. He's a Transformers guy. But even before that, at 12 years old, he was building robotics. He's much older than 12 now.
1
3
9
1,267
.@LeRobotHF is becoming the dominant open format for robot data. We recently introduced a native LeRobot reader in Daft, and we've now made it up to 15× faster.
2
5
16
2,086
We're excited to share that the Eventual team will be at AUTONOMOUS: The Future of Robotics & Physical AI We're looking forward to meeting with some of the most talented builders and investors in the industry. See you on July 16th at The Midway, San Francisco. #PhysicalAI #Robotics #AUTONOMOUS2026
3
309
When contributors roll through ... 🔥🔥🔥 🦾 @_BabTuna_ has been crushing PRs lately on daft so we invited him out to the office for lunch. Come back any time!
2
1
5
821
We trained this robot 🤖 to dance Bhangra 🕺🕺🕺through @UFBots Looks sooooooooo cool! #UltimateBotsStudio
3
299
As execution gets cheaper, thought leadership becomes even more important.
1
122
On a TPC-H repartition across 32 workers: at 10 TB the object-store shuffle runs the head node out of memory past 1,000 partitions, while Flight Shuffle completes every partition count and runs 3.6 to 4.7x faster where both finish.
1
1
1
114
Flight Shuffle makes disk the data plane. Each map task writes one combined Arrow IPC file to local disk. Every worker runs an Arrow Flight server that streams partitions over DoGet, so reduce tasks process batches as they land rather than buffering a whole partition first.
1
1
2
176
Daft's previous shuffle relied heavily on Ray's durable object store. Every map-to-reduce transfer became an object reference the head node tracked at ~3 KB each. Which means a 4096×4096 shuffle is 50 GB of references; and 8192×8192 is 200 GB. This would lead to head node OOMs.
1
3
4
289
If a distributed query has to materialize more than a few terabytes of data, there's one operation that will dominate: the shuffle. Shuffling data at scale has been a real bottleneck for Daft users, so we took the time to fix the root cause and rebuild the shuffle from scratch.
1
5
17
2,436
why did we choose arrow2... this ones for the real ones... universalmind303 what would we do without you 😩
1
4
868
Daft v0.7.4 just shipped. THE GREAT ARROW-RS MIGRATION 122 PRs completing the arrow2 → arrow-rs migration. 38,850 lines changed. Plus: full observability stack, Apache OpenDAL support, and Flight shuffle for Flotilla.
1
5
2,486
Agents do not fail because they lack intelligence. They fail because they lack institutional memory. In the new episode of Zero Shot Espresso with @akoratana, we discussed context graphs and decision traces, and what it takes for AI systems to actually improve over time in enterprise environments.
1
2
5
420
AI pipelines don't just run on Parquet files anymore. Audio, video, PDFs, source code: they're all central to modern workflows, but they're large, remote, and expensive to load eagerly. daft.File brings lazy, distributed file handling to all of them. One interface. Local or remote. Open only when needed. New blog walks through real examples: resampling audio, extracting video keyframes, parsing PDFs, and building code intelligence pipelines with AST, all inside Daft DataFrames. Link in the comments
1
1
2
144
The full episode is live if you're interested in piped.video/mgdJ90Sorvc?si=5WXZ…
1
181
AI value does not come from models alone. Sometimes it’s transformation. Sometimes it’s plumbing. Often it’s both. A clip from our latest Zero Shot Espresso episode with @frankzliu
1
3
175
When processing large amounts of data on GPUs, how do you maximize utilization so they're not just sitting idle waiting for data? Our engineer Srinivas Lade explains how we think about this problem at Daft. The key insight: a GPU is like a remote server sitting right next to you. When you send it work and wait for a response, you're stuck in a sequential bottleneck - process one batch, wait, get response, then send the next. How Daft solves this problem: piped.video/watch?v=WAsmZJ2k…
2
9
517
.@SourcetableApp CTO @andrewgrosser shares his recommended tech stack for serious startups - a "wicked combination" that includes: - S3 + Cassandra for data - Daft for processing - Python, WASM, Ray Learn how they built the first AI-powered spreadsheet: daft.ai/blog/how-sourcetable…
2
1
11
1,615
Fresh from NeurIPS! @EnricoShippole of @TeraflopAI processed 99% of US caselaw for under $1 using Daft and open sourced the data, the model, and the code. 7 million documents. 40 million pages. 365 years of case law. Here's how they went from 54GB compressed → 350GB CSV → cleaned Parquet, using distributed processing across bare-metal Kubernetes with Ray: daft.ai/blog/processing-99-o…
1
2
7
960
Feature Spotlight 💡 @mozilla Foundation's research noted that "Generative AI in its current form would probably not be possible without @CommonCrawl" - the open dataset of 250+ billion web pages that trained models like GPT-3. Daft makes accessing this massive dataset simple. What Daft provides: - Work seamlessly inside AWS (S3) or outside (HTTPS) - Choose your format: raw web archives (WARC), extracted text (WET), or metadata (WAT) - Start small (1 file for testing), scale up when ready Perfect for web-scale analytics or LLM training data preparation. docs.daft.ai/en/stable/datas…
2
311
English is now the new declarative language for data thanks to AI agents - you describe what you want over any type of data. When traditional SQL-based systems hit the limits of massive data, Spark was introduced. So what happens when AI systems & workflows hit the limits of massive multimodal data? Check out @JayChia5's argument in our new blog post: daft.ai/blog/agentic-systems…
1
3
284
Deep learning training & inference need highly optimized data loading to keep expensive GPUs busy - you want data ready when the GPU needs it. Machine Learning Engineer Hampus Londögård (@hlondogard) explored how Daft shines for this challenge: - Multi-modal namespace - Smooth Apache Arrow integration - Ray backend for petabyte-scale distributed workloads - Direct PyTorch integration via to_torch_iter_dataset() His mini benchmark vs native PyTorch DataLoaders: Daft held up really well, even surpassing PyTorch on full-size images. Full blog: blog.londogard.com/posts/202…
1
6
591
"For teams working primarily with structured data, what's the learning curve like for adopting Daft?" Kevin W., Daft's Founding Engineer says: If you're using SQL: Your existing SQL code works in Daft with minimal or no changes. If you're using Spark: We have Spark Connect support, so: - Daft can be a plug-in replacement for Spark - Works with PySpark and Spark SQL pipelines
1
4
214
That's a wrap for us at the Ray Summit! We chatted with 200+ people about their multimodal data workloads, batch inference needs, complex ETL jobs, and a lot more. We also gave out awesome prizes - a DGX Spark and two pink iPads to lucky winners. Thank you for stopping by!
1
6
408
LLM batch inference is often costly and slow - GPUs sit idle between batches and caches evict useful data. Today we're releasing a new inference backend in Daft that can cut batch inference time in half for workloads with common prefixes. We developed Dynamic Prefix Bucketing - a technique that intelligently routes and groups prompts to maximize cache usage while keeping GPUs continuously fed with work. The results on 200k prompts with shared prefixes across 128 GPUs: - 50.7% faster completion - Cache hit rate jumps from 29% → 54% Available today in Daft v0.6.9. Full technical deep dive: daft.ai/blog/cutting-llm-bat…
2
5
363
Want a chance to win Nvidia DGX Spark - the $4500 AI super computer? Come visit us at Ray Summit tomorrow!
3
2
24
5,394
Is Daft battle-tested vs alternatives like Spark and Ray Data? Spark: @essential_ai's @timr1126 processed 23.6B web documents (24T tokens) at 32,000 requests/second per VM: "We pushed Daft to the limit and it's battle tested... If we had to do the same thing in Spark, we would have to have the JVM installed, go through all of its nuts and bolts just to get something running." Ray Data: @CloudKitchens built their ML infrastructure on what they call the "DREAM stack" (Daft, Ray, poEtry, Argo, Metaflow): "[Daft] fills the gap of Ray Data by providing amazing DataFrame APIs" and "in our tests, it's faster than Spark and uses fewer resources." Additionally, independent testing by a ByteDance engineer running image classification on 1.28 million ImageNet images found that "Daft was about 20% faster than Ray Data."
2
3
334
A use case spotlight 💡 @CloudKitchens shared how they revamped their ML infrastructure, creating what they call the DREAM stack: - Daft as their "latest and greatest kitchen gadget" for ETL, analytics, and AI/ML at scale - Ray as a distributed compute engine for data and AI workloads - poEtry for Python dependency management - Argo as their trusted workflow orchestrator - Metaflow for a superior developer experience The results speak for themselves: - New data scientists can spin up projects in hours instead of weeks - All in Python - no more YAML struggles - Iteration speed increased several fold techblog.cloudkitchens.com/p…
6
497
GitHub's Dev Advocate @peckjon built an AI image search engine with Daft + @GitHubCopilot 🪄 What he liked about Daft: ✨ Native image processing - treats images as actual data, not file paths ⚡ Chainable operations: decode() → resize() → encode() 📊 SQL queries on non-relational, heterogeneous data "It's like the data science world finally caught up to what we dreamed of." Full article: peckjon.github.io/daft-image…
2
6
579
When to use Daft for batch inference? ✅ You need to run models over your data: Express inference on a column (ie. llm_generate, embed_text, embed_image) and let Daft handle batching, concurrency, and back-pressure ✅ You have data that are large objects in cloud storage: Daft has record-setting performance when reading from and writing to S3, and provides fe ✅ You're working with multimodal data: Daft supports datatypes like images and videos, and supports the ability to define custom data sources and sinks and custom functions over this data. ✅ You want end-to-end pipelines where data sizes expand and shrink: For example, downloading images from URLs, decoding them, then embedding them; Daft streams across stages to keep memory well-behaved. How does it work? Daft provides first-class APIs for model inference. Under the hood, Daft pipelines data operations so that reading, inference, and writing overlap automatically, and is optimized for throughput.
1
1
3
267
Full story in 3 mins 🎥: piped.video/watch?v=O8tDLssU…
2
152
Spark + 100MB PDFs = 💥 Ray Data + video pipelines = 😬 We built Flotilla so multimodal data *actually scales*: ⚡ Micro-batching ⚡ Streaming execution ⚡ Arrow Flight shuffle Benchmarks: 2–7× faster than Ray Data, 4–18× faster than Spark. #Daft #Multimodal #DataEngineering #OpenSource
5
4
227
In benchmarks, Flotilla was 2-7x faster than Ray Data and 4-18x faster than Spark. It was also the only engine that reliably finished multimodal workloads. daft.ai/blog/benchmarks-for-…
1
1
233
With Flotilla, we rethought the architecture: ⚡ Micro-batching + streaming → prevent OOMs ⚡ Concurrent pipelines → GPU, I/O, and writes run in parallel ⚡ Arrow Flight shuffle → 10× faster than Ray’s object store
1
1
153
Our Founding Engineer Kevin Wang will be speaking at next week’s Iceberg Meetup in San Francisco on Wednesday, October 1! If you have documents, images, and URLs in your Iceberg catalogs that you’re not sure how to make use of, then this is the talk for you! Kevin will share how you can leverage your existing Iceberg catalog to build scalable multimodal and AI applications powered by Daft. What you'll learn: - How to store your unstructured data together with your tabular data to optimize costs and performance - How Daft makes working with and exploring data of all modalities easy and performant - How you can go from raw data on Iceberg to generating embeddings with LLMs to uploading to your vectorDB all in one Daft query Kevin’s session is next Wed 10/1 at 5pm. Come check it out if you want to enable multimodal workloads with Iceberg!
1
2
174
Tired of staring at unhelpful progress bars while your data pipeline crawls along? We just published a 3-minute video on debugging performance bottlenecks in multimodal AI workloads. Start with a 70-second object detection script that offered zero visibility and end with a runtime cut by 65%. How? 3 steps: 1️⃣ Break down large UDFs 2️⃣ Use native Daft expressions 3️⃣ Tune UDF parameters. Give Daft a star (na2.hubs.ly/H01htFk0) and join our growing Slack community (na2.hubs.ly/H01hvBS0)! #Daft #Multimodal #DataEngineering #OpenSource
1
2
233
Thank you to everyone who joined our meetup last week! It was incredible to see so many developers facing the same multimodal data challenges we're solving with Daft. The energy and engagement from the community made it clear we're building something that truly matters. Special thanks to @ApertureData for hosting this event with us, and to our fantastic speakers @JayChia5, @vishakha041, @timr1126, and @vai_viswanathan. @DevSodhi did an excellent job moderating the panel discussion. Recordings of all talks and the panel will be available in a couple weeks - we'll share them here and in our community channels (Join our Slack! na2.hubs.ly/H01gvhZ0). Want to stay connected? Subscribe to our Luma calendar for upcoming events: na2.hubs.ly/H01gsTz0 And if you're excited about what we're building, we'd love a star on GitHub: na2.hubs.ly/H01gv-T0 Looking forward to seeing more of you at future events as we continue building the infrastructure for multimodal data processing together. #Daft #Multimodal #DataEngineering #OpenSource
2
6
289
Join us for this month's contributor sync where we'll showcase the latest developments transforming how developers work with multimodal data at scale. Here’s a sneak peak at the agenda! We've shipped major improvements across the stack: Model APIs and LLM operations with OpenAI integration for AI-powered pipelines, significant UDF performance improvements, and the new Turbopuffer integration for vector search workflows. Daft continues to evolve with our new .into_batches DataFrame API for efficient batch processing, Clickhouse data sink support for real-time analytics, Lance integration for modern columnar storage, and the new daft.File datatype that simplifies file handling across any storage backend. Our monthly deep dive features Colin Ho exploring Flotilla, Daft's distributed execution engine that powers these capabilities. We'll walk through the architecture that enables Daft to process petabytes of multimodal data—from the query planner to the Rust-powered execution layer. Colin will share benchmark results demonstrating how Flotilla handles real-world multimodal workloads: audio transcription at scale, video object detection across millions of frames, distributed image classification, and document embedding pipelines processing billions of tokens. Whether you're contributing code or building applications on Daft, you'll gain insights into the engine that makes "any modality, any scale" a reality. This Thursday 9/25 at 4pm PT. Add this event to your calendar: na2.hubs.ly/H01flVR0 #Daft #Multimodal #DataEngineering #OpenSource
2
112
Use case spotlight 💡 When John Shelburne (@thecatfix) needed to process 59 million bond market records (28M bonds + 31M prices + trade ideas) for his neural network model, he had a few options: - Continue using pandas (taking 3 days for processing) - Move to Vertex AI or another cloud service (more costs) Essentially, he was choosing between painfully slow local processing or expensive cloud compute. That's when he discovered Daft. By switching from pandas to Daft, John achieved: - 7.3 minutes total runtime (vs 3 days previously) - 600x faster processing - 50% lower memory usage - All running locally on his iMac As John put it: "HOLY SMOKES WAS IT FAST! My neural network model is now trainable in real time." If you're processing large-scale data and tired of waiting hours or days for your job to finish, check out Daft on GitHub: na2.hubs.ly/H01bZQy0
2
1
5
238
66% of enterprise data sits unused because traditional tools can't handle multimodal workloads. This Thursday 9/18 in SF: "Data Systems for AI: Multimodal Madness" Why existing infrastructure crumbles with multimodal data → How to run models on data → Real solutions from @ApertureData @essential_ai @voxel_labs #Daft #Multimodal #DataEngineering #OpenSource
2
1
2
355
Guess who’s back, back again? We’re bringing back monthly Daft’s contributor syncs! Last Thursday of every month. Our contributor community has been growing so much and we want to bring everyone together again. We’ll be giving updates about open source features and releases, Colin will be deep diving into Flotilla (our distributed engine), and we’ll have time at the end for any questions or further discussions. We’re planning to rotate the time for our sync every other month to accommodate for various time zones. This month’s sync will be on Thursday, September 25 at 4pm PT! Bring your questions and come hang out with the Daft crew! #Daft #Multimodal #DataEngineering #OpenSource
1
176