FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesAirflowApacheBigQueryCloudETLPySparkPythonSparkSQL
About the role
Key responsibilities & impact- Develop and maintain assigned batch ETL/ELT jobs using SQL, Python, and PySpark
- Support data ingestion from databases, APIs, files, and Google Workspace sources
- Perform basic data cleansing, transformation, and source-to-target validation
- Assist with development and maintenance of streaming data pipelines using Google Cloud Pub/Sub and Datastream
- Implement and maintain data quality checks for missing, duplicate, invalid, or delayed data
- Monitor batch and streaming pipelines, data freshness, and pipeline health
- Review logs, troubleshoot routine issues, and escalate complex incidents
- Assist with approved reruns and data reprocessing
- Prepare datasets for downstream analytics and reporting
- Help maintain dashboards monitoring pipeline status and data quality
- Document data mappings, transformation logic, validation rules, and troubleshooting steps
- Use Git and participate in code reviews and testing before changes are released
Requirements
What you’ll need- Suitable for fresh graduates or candidates with up to two years of relevant experience
- Knowledge of SQL, Python, and PySpark
- Ability to develop and maintain batch ETL/ELT jobs
- Ability to support data ingestion from databases, APIs, files, and Google Workspace sources
- Understanding of data cleansing, transformation, and source-to-target validation
- Ability to assist with streaming data pipelines using Google Cloud Pub/Sub and Datastream
- Ability to implement and maintain data quality checks
- Ability to monitor pipelines, review logs, and troubleshoot routine issues
- Ability to document data mappings, transformation logic, validation rules, and troubleshooting steps
- Experience or familiarity with Git, code reviews, and testing
- Familiarity with Google Cloud services including BigQuery, Dataform, Google Cloud Storage, Managed Service for Apache Airflow, Managed Service for Apache Spark, Pub/Sub, and Datastream
Benefits
Comp & perks- Hands-on experience with the team’s Google Cloud data stack
- Guidance and code reviews from experienced team members
