Data Nerds! I ranked every data engineering tool by how often it shows up in 4M+ job postings. π
But here's the catch π³.
Some critical skills show up way less than they should because they're often assumed to be foundational skills for jobs. (e.g., Skills like Bash/Terminal for running pipelines)
Anyway, here's the breakdown of the tiers π (Note: % = how often each tool appears in DE job postings)
π΄ S TIER β Non-Negotiable
The core skills needed for any DE job. Don't apply without these:
π SQL (~68%) β every warehouse runs on it. Query, transform, and model data.
π Python (~67%) β the pipeline language. Ingestion, automation, APIs, glue between systems.
β¨οΈ Terminal/Bash (~11%) β every tool you'll use runs from here. This is highly undervalued in postings.
π Git (~11%) β version control. Every team uses it. Same posting-% caveat as Bash.
βοΈ One cloud platform + warehouse (~26-46%) β AWS + Redshift, GCP + BigQuery, or Azure + Synapse. Combined cloud presence is in nearly every posting.
Start with SQL, then Python. Everything else you absorb alongside them.
π A TIER β Job-Ready Foundation
The tool that closes the gap from "learning DE" to "hireable for modern stacks":
πͺ dbt (~10%) β only 10% of all DE postings, but 36% in Analytics Engineer (AE) roles.
That's not a niche, it's a leading indicator. AE is the new hybrid role modern data teams are hiring for: part analyst, part engineer.
β
Land the job with S + A. Pass the interview with conceptual knowledge of B Tier π
π‘ B TIER β Interview-Aware
Know what they solve. Don't expect to code from scratch:
βοΈ Airflow (~17%) β orchestration. Built on DAGs (directed acyclic graphs).
β‘ Spark (~38%) β distributed computing for processing large datasets.
π Kafka (~19%) β real-time event streaming between systems.
All these depend on a foundational knowledge of Python & SQL; don't jump the gun learning these.
π’ C TIER β Data Platform Awareness
Pick the one your company uses. Understand both conceptually:
βοΈ Snowflake (~26%) β pure SQL warehouse. Optimized for analytics. Modern-stack favorite.
π§± Databricks (~24%) β lakehouse on Spark. Handles structured + unstructured. ML/AI heavy teams.
π΅ D TIER β Versatility Multipliers
Lower headline demand, but high value per hour:
π Power BI (~15%) / Tableau (~10%) β but the kicker: in AE roles these jump to 28% / 33%.
Modern data teams want pipeline builders who can also visualize. For analysts pivoting to DE, lead with this in interviews.
π£ E TIER β Path-Dependent
High demand on paper, but concentrated in legacy enterprise stacks. Skip until your job requires it:
β Java (~25%) β legacy enterprise data infrastructure
βοΈ Scala (~22%) β Spark's native language. Spark-heavy shops.
π₯ How did I derive this ranking? In my latest video, I walk through the concepts first (the DE lifecycle, what each tool actually solves) and then derive the tiers. (Link in comments π)