🦸♂️ Not all data engineers wear cape, but they should.
Migrations are high-stakes, high-risk, & full of villains (👋 misaligned timelines).
Enter The Data Migration Universe for real stories, best practices & a readiness check. Are you migration-ready?
get.datafold.com/x-dms
Data-driven design… in cookie form! 🖤🍪 Our team's holiday baking activity was the perfect mix of innovation and fun with the Datafold logo as a 3D-printed cookie cutter!
🚀 New Datafold x Power BI integration!
Track data lineage &assess impact on Power BI assets with ease. Proactively prevent data quality issues before they reach your dashboards.
Learn more 👇
get.datafold.com/power-bi-bl…
Data migrations shouldn’t take years.
Meet the Datafold Migration Agent: AI-powered translation & validation that works across any SQL dialect, orchestration framework, or GUI platform. Plus, it refines until you reach full parity.
get.datafold.com/dma-blog-tw…
Monitors are a powerful way to help you automatically track and manage data quality across your stack.
From schema changes to anomaly detection, Datafold’s Monitors catch issues before they impact your business.
get.datafold.com/monitors-tw…
We’ve heard the challenges data teams face—ensuring data integrity amidst scaling pipelines & growing complexity.
Today, we’re taking a big step to address those pains.
Introducing Monitors in Datafold: Data monitoring that starts upstream.
get.datafold.com/monitors-bl…
⚡ Introducing No-Code CI in Datafold ⚡
Easily integrate data diffing into your code reviews, regardless of your data tooling.
Use stored procedures, SQL models, Airflow, or other technologies to push data diff results directly in PR comments.
get.datafold.com/no-code-ci-…
The best kind of data quality testing is proactive.
It’s how all data engineering teams should operate, but largely don’t because there’s a lot of noise around what are the right data quality principles, tools, and workflows.
We know that in this day and age, data teams have to get the data *right*, and feel immense pressure to do so.
After all, it’s high data quality that protects the integrity of our executive dashboards, machine learning models, and reverse ETL syncs.
If you only use dbt tests for data quality, the bad news is that it prevents only some data quality issues. @leoebfolsom and @elliot_j_g discuss why your data team needs Datafold too for complete data quality test coverage:
datafold.com/blog/difference…
🎉 Big news! Datafold + @dremio are teaming up to supercharge your data workflows! 🎉
🚀 Speed up your move to a modern lakehouse with cross-database diffing
🪴 Enhance your dbt project’s data quality with CI testing
🔗 get.datafold.com/dremio-twit…
Our founder, @glebmm, was recently on the Data Engineering Podcast to discuss:
🧐 Why validating + reconciling data at scale remains so challenging
👯 Common patterns that data teams encounter
⏩The future of data engineering tooling
Read the summary 👇
datafold.com/blog/the-data-e…
Yesterday, we announced the ability to schedule data diffs across databases for ongoing data replication testing.
📽️ Watch @leoebfolsom demo how to use Datafold’s cross-database diffing + new Monitors for continuous source-to-target validation.
get.datafold.com/l/1011581/2…
With Datafold, test your replication pipelines automatically and continuously; validate consistency between source and target at any scale, and receive alerts about any discrepancies.
get.datafold.com/replication…
We know how important it is for the data you’re replicating across databases to be right. We know this data is often mission-critical—powering core analytics work and machine learning models, and guaranteeing data reliability and accessibility is vital.
As data teams scale and mature, they don't always know about best practices that can build more efficient, error-resistant, and collaborative CI pipelines.
We're here to fix that with our latest tutorial on adding 4 advanced CI integrations for your dbt project's CI workflow.
👨👧👧 Representative sampling: Establish difference "thresholds" to stop diffs once the set number of differences has been found per column, further reducing time (and costs) spent during validation.