Apache DataFusion is becoming a foundational building block for modern data systems.
From RisingWave, InfluxDB, LanceDB, and GreptimeDB to OpenObserve, ParadeDB,
Spice.ai, Vortex, Apache Comet, Ballista, and many others, a growing ecosystem is building on the same shared query foundation.
Why?
Building a high-performance query engine from scratch requires a huge amount of engineering.
A modern engine needs:
SQL and DataFrame APIs
Logical and physical query planning
Query optimization
Vectorized execution
Multithreaded processing
Streaming execution
File-format support
Object-storage integration
Memory management
Extensible functions and operators
Apache DataFusion provides these core capabilities in a modular and extensible engine written in Rust and built around the Apache Arrow in-memory format.
The real value is not just performance.
It is extensibility.
This allows teams to reuse the common foundations of query processing without giving up control over the parts that make their systems unique.
Instead of rebuilding the same SQL parser, optimizer, execution engine, and file-format support again and again, projects can focus their engineering effort on their specific workloads and product differentiation.
RisingWave follows the same approach.
RisingWave uses Apache DataFusion as its default batch query engine while continuing to use its streaming-native engine for incremental and stateful stream processing.
Apache DataFusion handles the query engine, allowing projects to focus on the unique requirements of their specific use cases.