Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Altarum

Data Engineer – Generative AI, Healthcare Data

Altarum

. Design and operate cloud data platform components, including lakehouse storage, compute, SQL and ELT engines, and reusable transformation workflows using Databricks or comparable platforms .

Posted 9/29/2026full-timeUnited StatesMid-LevelSenior💰 $144,771 - $174,723 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and operating cloud data platforms, with strong capabilities in Python and SQL for building production systems and generative AI applications. Proficient in implementing CI/CD processes, data quality management, and compliance with healthcare data standards.

Highest-signal resume keywords
Python ProgrammingSQL DevelopmentDatabricks PlatformGenerative AI Application DevelopmentCI/CD Implementation

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Data EngineeringCloud EngineeringAutomated TestingData TransformationInfrastructure as CodeDistributed ProcessingData ModelingData Quality ManagementAPI IntegrationVersion Control
Soft Skills
Technical LeadershipMentoringClear CommunicationProblem SolvingCollaboration
Tools & Technologies
DatabricksSparkMLflowLangChainLlamaIndexDelta LakeDbtGitAWSAzure
Certifications & Qualifications
Databricks Certified Generative AI Engineer Associate
Industry Keywords
Healthcare DataMedicaidFHIRHL7Data GovernanceRegulated ReportingPrivacy by DesignNIST AI Risk Management FrameworkQuality MeasuresDocument Processing

Tech Stack

Tools & technologies
AWSAzureCloudPySparkPythonSparkSQLUnity

About the role

Key responsibilities & impact
  • Design and operate cloud data platform components, including lakehouse storage, compute, SQL and ELT engines, and reusable transformation workflows using Databricks or comparable platforms
  • Develop production Python and SQL code, versioned reusable packages, APIs, and configuration-driven frameworks with automated tests and documentation
  • Own CI/CD for data pipelines, AI applications, and supporting services, including environment configuration, automated validation, deployment, release verification, and rollback
  • Implement infrastructure as code and secure deployment patterns in AWS, AWS GovCloud, Azure, Azure Government, or other approved environments
  • Own ingestion, transformation, and normalization pipelines for structured, semi-structured, and unstructured healthcare and public health data, including batch and streaming workloads
  • Build reusable connectors, data contracts, and curated datasets using Python, SQL, Spark or PySpark, Delta Lake, dbt, or comparable technologies
  • Integrate Medicaid and Medicare claims and encounters, FHIR R4, US Core, HL7 v2, social determinants of health, geospatial data, surveys, and qualitative sources as projects require
  • Translate analytics, reporting, evaluation, and data science requirements into tested data models and dependable datasets with documented lineage
  • Define and implement data quality rules, validation, lineage, metadata, and monitoring
  • Diagnose and resolve production failures, data quality issues, and performance bottlenecks; implement root-cause fixes and maintain operational runbooks
  • Establish reliability and freshness targets; own monitoring, incident response, and recovery for assigned services
  • Optimize compute, storage, partitioning, and pipeline execution; track cost and performance and implement guardrails
  • Own AI application delivery from requirements and solution design through implementation, evaluation, deployment, and production support
  • Build retrieval-augmented generation and document-intelligence applications, including ingestion, chunking, embeddings, vector search, and retrieval
  • Use MLflow or comparable tools to manage generative AI evaluation and tracing, evaluation datasets, custom scorers, and prompt/model versioning
  • Develop LLM applications and agentic workflows using Python, LangChain, LlamaIndex, or comparable frameworks
  • Deploy and operate AI applications and model endpoints using Databricks Model Serving or comparable services, with governed data access, monitoring, and CI/CD
  • Build reusable AI services and automation for healthcare reporting, document extraction, and knowledge retrieval
  • Apply privacy by design, encryption, least-privilege access, secrets management, and secure networking
  • Implement technical controls and documentation for HIPAA, 42 CFR Part 2, IRB and data-use agreement restrictions, and responsible AI practices informed by the NIST AI Risk Management Framework
  • Document model behavior, evaluate fairness and explainability, and maintain model cards and governance artifacts
  • Serve as technical lead for defined projects or workstreams; coordinate contributors, review implementation choices, and resolve blockers
  • Mentor colleagues through pairing, code reviews, and technical guidance; lead workshops and create tutorials, runbooks, and examples
  • Explain technical concepts and tradeoffs to varied audiences and contribute technical materials for proposals and client engagements

Requirements

What you’ll need
  • Bachelor's degree in computer science, data science, information systems, engineering, mathematics, statistics, public health informatics, or a related field, or equivalent practical experience
  • Typically 3–5 years of relevant experience in data engineering, AI engineering, software engineering, analytics engineering, cloud engineering, or applied technical research, including independent delivery of production systems
  • Strong Python and SQL skills, with experience building maintainable code, data models, automated tests, and reliable ingestion/transformation pipelines
  • Hands-on experience with a cloud platform and a modern data platform such as Databricks or an equivalent, including distributed processing with Spark or comparable technology
  • Experience using Git, CI/CD, environment-specific configuration, and automated deployment to deliver and support production workloads
  • Hands-on experience building and deploying generative AI applications with LangChain, LlamaIndex, or comparable frameworks, including retrieval-augmented generation or document processing, prompt design, API integration, evaluation, and monitoring
  • Ability to turn requirements into technical designs, manage defined work, troubleshoot independently, and guide collaborators through clear explanations and constructive feedback
  • Databricks Certified Generative AI Engineer Associate certification strongly preferred; comparable generative AI credentials with relevant practical experience also valued
  • Databricks Mosaic AI, Vector Search, Model Serving, MLflow evaluation and tracing, and Unity Catalog
  • Experience delivering governed generative AI applications
  • Healthcare and Medicaid data, quality measures, regulated reporting, FHIR or HL7, or accessible document generation
  • Databricks workflows, Asset Bundles, Delta Lake, PySpark, dbt, orchestration, or infrastructure as code
  • Experience leading small technical teams, mentoring colleagues or students, teaching technical subjects, or supervising applied research
  • Candidates must be currently eligible to work in the United States; sponsorship is not available
  • All work must be performed within the continental U.S. for the duration of employment, unless required by contract
  • Ability to work core hours aligned with Eastern Time, unless otherwise approved by your manager
  • Remote employees must maintain a dedicated, ergonomically appropriate workspace free from distractions, with reliable internet and a mobile device that supports efficient work
  • Travel Requirements: <10%

Benefits

Comp & perks
  • Competitive Medical, Dental and Optical plans
  • Generous Paid Time Off
  • 8 Company observed holidays plus 3 floating holidays
  • Tuition Assistance
  • 401K Plan (3% employer contribution plus opportunity for gainsharing)
  • Life, AD&D & Disability coverage
  • Flexible work environment
  • Remote work with occasional in-person collaboration days
  • Virtual participation in Collaboration Days for non-local employees