FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and maintaining production data pipelines, with strong capabilities in Python development, data modeling, and performance tuning. Proficient in implementing data quality checks and collaborating with cross-functional teams to ensure data integrity and compliance.
Highest-signal resume keywords
Python DevelopmentData Pipeline EngineeringSQL Performance TuningAzure Data FactoryData Modeling Fundamentals
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Data Pipeline DesignPySpark DataFrameSQL APIsData Quality ChecksVersion ControlInfrastructure as CodeStreaming Data ProcessingData ModelingRoot-Cause AnalysisCI/CD Practices
Soft Skills
Excellent CommunicationAnalytical Problem-SolvingConsultative Technical SkillsProject Management
Tools & Technologies
Azure Data FactoryDatabricksSynapse/FabricADLSTerraformDockerKubernetesKafkaEvent HubsDelta Lake
Industry Keywords
HIPAASOC 2PCIGDPR
Tech Stack
Tools & technologiesAzureDockerKafkaKubernetesPySparkPythonSparkSQLTerraform
About the role
Key responsibilities & impact- Design, build, and maintain idempotent, observable, and recoverable batch and streaming data pipelines
- Model data for analytics, including dimensional models, semantic layers, and curated marts
- Integrate data from operational databases, SaaS APIs, files, and event streams while handling schema drift and late-arriving data
- Build pipeline-integrated data quality checks for freshness, volume, uniqueness, and referential integrity
- Define failure alerting and escalation processes
- Own production pipelines through monitoring, on-call data incident response, root-cause analysis, and backfills
- Tune performance and cost through partitioning, clustering, file sizing, and warehouse or cluster sizing
- Apply version control, code review, CI/CD, automated testing, and infrastructure as code
- Implement access controls, PII handling, retention, lineage, and audit requirements with security and compliance teams
- Partner with analysts, data scientists, and product engineers to create durable data contracts
- Maintain data dictionaries, lineage, and pipeline runbooks
Requirements
What you’ll need- 5+ years building production data pipelines
- Strong hands-on Python development for data engineering, including testing, packaging, and code review practice
- Working knowledge of PySpark DataFrame and SQL APIs, joins and aggregations at scale, partitioning and shuffle behavior, and Spark UI diagnosis
- Strong SQL skills, including window functions, query plans, and performance tuning
- Hands-on experience with Azure Data Factory, Databricks, Synapse/Fabric, and ADLS
- Solid data modeling fundamentals, including normalization, star schemas, and slowly changing dimensions
- Git-based workflow and experience shipping through CI/CD
- Excellent communication skills, including explaining pipeline issues to non-technical stakeholders
- Exceptional analytical and problem-solving skills, including root-cause analysis
- Strong knowledge and experience working with customers in a consultative technical environment
- Streaming experience with Kafka or Event Hubs
- Experience with Delta Lake or Iceberg
- Infrastructure as code with Terraform or Bicep and containerization with Docker or Kubernetes
- Experience in regulated environments involving HIPAA, SOC 2, PCI, or GDPR
- Experience building ML data platforms or supporting feature pipelines
- Ability to manage multiple client projects and deliver high-quality results on time
- Experience with Azure DevOps or GitHub for source control and pipelines
