When to use Daft for batch inference?
✅ You need to run models over your data: Express inference on a column (ie. llm_generate, embed_text, embed_image) and let Daft handle batching, concurrency, and back-pressure
✅ You have data that are large objects in cloud storage: Daft has record-setting performance when reading from and writing to S3, and provides fe
✅ You're working with multimodal data: Daft supports datatypes like images and videos, and supports the ability to define custom data sources and sinks and custom functions over this data.
✅ You want end-to-end pipelines where data sizes expand and shrink: For example, downloading images from URLs, decoding them, then embedding them; Daft streams across stages to keep memory well-behaved.
How does it work?
Daft provides first-class APIs for model inference. Under the hood, Daft pipelines data operations so that reading, inference, and writing overlap automatically, and is optimized for throughput.