FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesAirflowApacheAWSBigQueryCloudDistributed SystemsDockerGoogle Cloud PlatformHadoopHDFSKafkaKubernetesPySparkScalaSparkSQLTerraformYarn
About the role
Key responsibilities & impact- Lead migrations of enterprise Spark workloads from on-premise environments to AWS and GCP
- Assess Spark applications, clusters, configurations, dependencies, data flows, and resource utilization
- Determine migration approaches across rehost, replatform, refactor, modernize, or retire
- Modernize traditional cluster-based workloads for serverless Spark where appropriate
- Design and implement architectures using AWS EMR Serverless, S3, Glue, Lake Formation, GCP Dataproc Serverless, GCS, and BigQuery
- Refactor legacy PySpark/Scala/Spark SQL applications for cloud portability, scalability, and reliability
- Migrate workloads using Hadoop, HDFS, YARN, Hive, and on-premise Spark clusters
- Troubleshoot and optimize Spark workloads, including partitioning, shuffle behavior, joins, data skew, execution plans, executor configuration, serialization, and SQL execution
- Benchmark performance and optimize serverless workloads for performance, reliability, and cloud cost
- Build reusable migration tooling, automation, templates, and frameworks
- Implement CI/CD and Infrastructure as Code using tools such as Terraform
- Define testing, validation, cutover, rollback, observability, and production-readiness patterns
- Partner with Data Engineering, ML, Cloud Architecture, Platform Engineering, DevOps/SRE, Security, Governance, and FinOps teams
- Own the complete migration lifecycle: Discover, Assess, Design, Refactor, Migrate, Validate, Optimize, Operate
Requirements
What you’ll need- 8+ years of experience across data engineering, distributed systems, cloud engineering, or platform engineering
- 5+ years of hands-on Apache Spark experience in enterprise environments
- Strong PySpark and/or Scala development experience
- Proven experience migrating large-scale Spark workloads between infrastructure platforms
- Hands-on experience with both AWS and GCP
- Experience with on-premise Hadoop/Spark ecosystems, including HDFS, YARN, and Hive
- Deep understanding of Spark internals and distributed processing
- Strong SQL and data engineering fundamentals
- Experience with cloud data lakes and object storage
- Strong production troubleshooting and performance-tuning experience
- Experience with CI/CD, Git, and Infrastructure as Code
- Ability to own migration work end-to-end, from discovery and architecture through production cutover and optimization
- Experience with EMR/EMR Serverless, Dataproc/Dataproc Serverless, Glue, Lake Formation, BigQuery, Delta Lake, Iceberg, Kafka, Airflow, Terraform, Docker, or Kubernetes is valuable
Benefits
Comp & perks- Remote work arrangement
- Coverage during Pacific Hours (8:00 AM–5:00 PM PST)
