you bring to the team? 5+ years of data engineering experience, building and operating production data pipelines at scale (TB+ datasets, hourly/daily batch or streaming workloads). Hands-on production experience with Apache Spark and distributed data processing frameworks such as Flink, Hive, or Trino. Strong understanding of large-scale batch and streaming pipelines, including performance tuning and troubleshooting. Language is not a filter: Scala, Python, or Java are all fine. What matters is that you can debug and ship production Spark code, not which language you write it in Production experience building and operating data solutions on GCP or AWS, including cloud-native services such as BigQuery, Dataproc, GCS, S3, EMR, or Redshift. Experience across the full project lifecycle is