Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Rezilient Health

Senior Platform Engineer

Rezilient Health

. Design, provision, and maintain cloud infrastructure using Infrastructure as Code, managing development, staging, and production environments with repeatable, auditable configuration.

Posted 10/8/2026full-timeRemote • United StatesSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and maintaining cloud infrastructure using Infrastructure as Code, with a strong focus on CI/CD pipeline management, observability, and incident response. Proficient in containerization and orchestration, ensuring scalability and security in high-availability environments.

Highest-signal resume keywords
Cloud Infrastructure ManagementCI/CD Pipeline DevelopmentContainerization and OrchestrationObservability and MonitoringIncident Response Leadership

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Infrastructure as CodeTerraformDockerKubernetesCI/CD PipelinesPythonGoBashGitCloud Platforms
Soft Skills
Excellent CommunicationDetail-OrientedCalm Under PressureMethodical ApproachComfortable with Ambiguity
Tools & Technologies
DatadogPrometheusGrafanaCloudWatchELK/OpenSearchJiraConfluence
Industry Keywords
Site Reliability EngineeringDevOpsHealthcareHIPAAData-Sensitive Environments

Tech Stack

Tools & technologies
AWSAzureCloudDockerGoogle Cloud PlatformGrafanaKubernetesPrometheusPythonTerraformGo

About the role

Key responsibilities & impact
  • Design, provision, and maintain cloud infrastructure using Infrastructure as Code, managing development, staging, and production environments with repeatable, auditable configuration.
  • Own and evolve CI/CD pipelines with automated testing gates, blue/green and canary rollouts, and reliable rollbacks.
  • Build and operate observability across metrics, logging, distributed tracing, dashboards, and alerting; define SLOs, SLIs, and error budgets.
  • Lead incident response, including on-call participation, triage, mitigation, communication, and blameless postmortems.
  • Automate operational toil through scripting and tooling to improve detection and recovery times.
  • Design for scalability and resilience through capacity planning, load and failure testing, autoscaling, redundancy, disaster recovery, and backup strategies.
  • Manage Docker and Kubernetes workloads, including networking, service discovery, and resource management.
  • Partner with Security and Engineering to harden infrastructure through secrets management, network segmentation, encryption, vulnerability scanning, patching, and audit logging consistent with HIPAA-aligned PHI requirements.
  • Collaborate with development teams to embed reliability and operability into services from design through production.

Requirements

What you’ll need
  • Bachelor's degree in computer science, software engineering, or a related field, or equivalent hands-on experience.
  • 5+ years in Site Reliability, DevOps, or infrastructure engineering roles at startup or growth-stage organizations, ideally within healthcare, health tech, or another regulated, high-availability, data-sensitive industry.
  • Deep hands-on experience operating production systems on a major cloud platform (AWS, GCP, or Azure), including compute, networking, storage, and managed database services.
  • Strong proficiency with Infrastructure as Code (Terraform or equivalent) and configuration management.
  • Production experience with containerization and orchestration (Docker and Kubernetes), including deployment, scaling, and troubleshooting.
  • Expertise building CI/CD pipelines and release automation, with a track record of enabling safe, frequent deployments.
  • Hands-on experience with observability and monitoring tooling (e.g., Datadog, Prometheus, Grafana, CloudWatch, ELK/OpenSearch) and defining SLOs, SLIs, and error budgets.
  • Strong scripting and automation skills (Python, Go, or Bash) and fluency with command-line and cloud CLIs.
  • Experience leading incident response and on-call, including postmortem and root-cause analysis practices.
  • Solid understanding of infrastructure and network security, secrets management, and encryption, with familiarity handling sensitive or protected data (PHI/HIPAA experience strongly preferred).
  • Proficient with Git and version control workflows; familiarity with the Agile Development Framework and ideally the Atlassian toolset (Jira and Confluence).
  • Excellent verbal and written communication skills; calm, methodical, and detail-oriented under pressure, and comfortable owning ambiguity in an early-stage environment.

Benefits

Comp & perks
  • Meaningful ownership through stock options, allowing you to share directly in the value you help create as we grow
  • Competitive base compensation
  • Comprehensive medical, dental, and vision coverage options, with Rezilient contributing toward your premiums
  • Complimentary access to Rezilient's clinical programs for you and your household members, the same connected care we deliver to our patients
  • 401(k) retirement plan to support your long-term financial well-being
  • Flexible Paid Time Off, so you can recharge without counting days
  • Dedicated Paid Sick Leave, separate from your FTO, for when health needs come first
  • 11 paid company holidays each year
  • Paid family leave to support you through life's biggest moments
  • Optional ancillary benefits, including life insurance, disability coverage, and a Health Savings Account (HSA)