FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and managing large scale Apache Spark data processing jobs, with a focus on performance tuning and data modelling for analytics. Proficient in automated scheduling, monitoring, and ensuring data quality for efficient data pipeline execution.
Highest-signal resume keywords
Apache Spark Performance TuningData Modelling for AnalyticsAutomated Scheduling and MonitoringApache Iceberg or Open Table FormatsTrino or Similar Query Engines
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Apache SparkData ModellingPerformance TuningAutomated SchedulingData Quality Checks
Tools & Technologies
Apache IcebergTrinoLakehouse Architecture
Industry Keywords
Batch Data PipelinesData Processing JobsEnd to End Data FlowsOn-Premises Cluster Experience
Tech Stack
Tools & technologiesApacheSpark
About the role
Key responsibilities & impact- Build and run large scale batch data pipelines.
- Keep data processing jobs fast and affordable as data volumes grow.
- Build and manage large scale Apache Spark data processing jobs.
- Perform performance tuning to maintain efficient and reliable pipeline execution.
- Support data modelling for analytics.
- Organise data in a lakehouse for downstream use.
- Handle automated scheduling, monitoring, and data quality checks.
- Work with platform and product teams on end to end data flows.
Requirements
What you’ll need- Strong hands on Apache Spark including performance tuning, not Spark usage through a managed notebook only.
- Data modelling for analytics and organising data in a lakehouse so it is usable downstream.
- Automated scheduling, monitoring, and data quality checks.
- Works with platform and product teams on end to end data flows.
- Apache Iceberg or other open table formats.
- Trino or similar query engines.
- On premises or self managed cluster experience.