FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Data Engineer – Legacy Systems, AI Workflows
Groundswell. Design, develop, and operate batch and streaming ETL/ELT pipelines that ingest data from multiple structured and semi-structured sources into secure cloud data platforms .
Tech Stack
Tools & technologiesAWSCloudETLJavaNumpyPandasPythonSOAPSQL
About the role
Key responsibilities & impact- Design, develop, and operate batch and streaming ETL/ELT pipelines that ingest data from multiple structured and semi-structured sources into secure cloud data platforms
- Onboard legacy data sources by reviewing schemas, mappings, interfaces, validation rules, ownership, and operational support procedures
- Build automated data-quality checks for completeness, accuracy, consistency, timeliness, and referential integrity, with actionable monitoring and alerting
- Automate deployments and pipeline operations using SOAP APIs, workflow orchestration, and environment-specific configurations
- Troubleshoot failures across distributed data systems using logs, metrics, lineage, and operational signals; communicate root causes and recovery plans clearly
- Collaborate with software engineers, cloud engineers, data scientists, security teams, architects, and customer stakeholders to translate mission needs into measurable data products
Requirements
What you’ll need- Bachelor’s degree in Computer Science, Computer Engineering, Mathematics, Statistics, or a related technical field
- 5+ years of professional data engineering experience, including production pipeline development and operations
- Proficiency in Python, Java, R, or other programming language used for data engineering
- Experience with legacy systems that utilize SOAP APIs for data interaction
- Experience implementing data-quality frameworks, validation processes, monitoring, and alerting for production pipelines
- Working knowledge of data security, access controls, encryption, auditability, and governance in a regulated, restricted, or compliance-oriented environment
- Experience working with APIs (SOAP/REST) and structured data formats such as JSON, XML, CSV, and Parquet
- Ability to document technical decisions and collaborate with multidisciplinary teams and customer stakeholders
- Must be a U.S. Citizen per contract requirements
- Must be able to obtain and maintain a Public Trust Clearance in accordance with contract requirements
- Preferred: experience building preprocessing, chunking, filtering, metadata-enrichment, or evaluation pipelines for LLM and other AI/ML workflows
- Preferred: experience operating data platforms with segmented networks, limited connectivity, strict change control, or formal authorization requirements
- Preferred: strong SQL expertise, including query optimization, data modeling, joins, window functions, and analysis of large datasets
- Preferred: experience with Jupyter Notebook or equivalent tools for analyzing data
- Preferred: experience with Python libraries such as numpy and pandas
- Preferred: hands-on experience with at least one major cloud provider; AWS strongly preferred
- Preferred: C3.ai experience
- Active Public Trust Clearance preferred
- Preference given to candidates local to the Washington, DC metro area
- AWS Data Engineer Certification, Databricks, cloud data engineering, or other relevant professional certification preferred
Benefits
Comp & perks- Comprehensive medical, dental, and vision plans
- Flexible Spending Account
- 4% 401K Match (immediate vesting)
- Paid Time Off
- Tuition reimbursement, certification programs, and professional development
- Flexible work schedule
- On-site gym and childcare option