FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Data Engineer, AWS, Spark
OPENDataJobs. Build Spark-based extract, transform, and load (ETL) pipelines with Glue, Amazon EMR, Lambda, and Step Functions .
Posted 9/17/2026full-timeWashington • District of Columbia • United StatesMid-LevelSenior💰 $103,000 - $140,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building Spark-based ETL pipelines on AWS, utilizing tools such as Glue, Amazon EMR, and Lambda. Proficient in data quality, validation, and lineage, with a strong focus on designing scalable data storage solutions and automating data ingestion processes.
Highest-signal resume keywords
Spark ETL On AWSPython And PySparkS3 Data-Lake DesignData Quality And ValidationInfrastructure-As-Code
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Data EngineeringETL PipelinesEvent OrchestrationData QualityValidationPythonPySparkS3Apache IcebergTrino
Soft Skills
CollaborationProblem SolvingCommunication
Tools & Technologies
AWS GlueAmazon EMRLambdaStep FunctionsCloudFormationTerraformPowerPointExcel
Certifications & Qualifications
Bachelor's DegreeMaster's Degree PreferredPublic Trust Determination
Industry Keywords
Federal Information TechnologyHigh-Volume DataLegacy ETL MigrationApache Ranger
Tech Stack
Tools & technologiesApacheAWSDynamoDBETLPostgresPySparkPythonSparkSQLTerraform
About the role
Key responsibilities & impact- Build Spark-based extract, transform, and load (ETL) pipelines with Glue, Amazon EMR, Lambda, and Step Functions
- Write processing in Python and PySpark
- Design the S3 layer, including Parquet, partitioning, and lifecycle, feeding Apache Iceberg tables
- Connect PostgreSQL on Amazon Aurora and DynamoDB, with Trino for federated SQL across them
- Build ingest, processing, and storage at scale
- Deliver clean and trustworthy data pipelines for high-volume federal platforms
- Automate ingestion previously handled case by case
- Build storage designs supporting downstream systems
- Work alongside developers, engineers, data scientists, architects, and strategists
- Participate in work ranging from strategy to implementation
- Adapt development workflows using AI tools to implement, test, and improve owned components
- Discuss a difficult problem, decisions made, outcomes, and lessons learned
Requirements
What you’ll need- 4+ years of data engineering experience
- Bachelor's degree
- Spark ETL on AWS (Glue, Amazon EMR) in Python and PySpark
- S3 data-lake design with Parquet, partitioning, and lifecycle feeding Apache Iceberg tables, Amazon Aurora PostgreSQL, and DynamoDB
- Event orchestration using Lambda, Step Functions, SQS/SNS, secrets management, and monitoring
- Data quality, validation, and lineage
- Infrastructure-as-code using CloudFormation or Terraform
- Basic proficiency in writing, PowerPoint, and Excel
- United States citizenship
- Ability to obtain a Public Trust determination
- Master's degree in a relevant field preferred
- Trino or comparable federated SQL preferred
- Apache Ranger-governed access preferred
- Legacy ETL migration, such as DataStage, preferred
- Federal information technology or high-volume data experience preferred
- Familiarity with AI-assisted developer tooling preferred
Benefits
Comp & perks- Medical, dental, and vision with the employee premium fully paid and half of dependent premiums
- Employer-paid life, accidental death, and short-term and long-term disability insurance
- 401(k) matched 100% up to 4% of salary, vesting immediately
- Unlimited paid time off
- Sponsored professional certifications and continuing education
- Extensive onboarding
- Rotation across functions to expand skills and perspective