Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Peraton

Lead Site Reliability Engineer

Peraton

. Lead infrastructure-level disaster recovery drill execution .

Posted 9/29/2026full-timeUnited StatesSenior💰 $112,000 - $179,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in disaster recovery processes, infrastructure automation, and cloud engineering, with a strong focus on AWS services and Infrastructure as Code. Proficient in managing Kubernetes clusters and developing observability solutions to ensure system reliability and compliance.

Highest-signal resume keywords
Disaster Recovery ValidationInfrastructure as Code (IaC) Using TerraformKubernetes AdministrationAWS Services ExpertiseCI/CD Pipeline Development

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Disaster Recovery ProcessesInfrastructure AutomationKubernetes ManagementAWS CloudFormationPython ProgrammingJava ProgrammingCI/CD Deployment PipelinesData Integrity ChecksMonitoring Tools DevelopmentContainer Orchestration
Soft Skills
Analytical SkillsDocumentation SkillsCollaboration SkillsProblem-Solving SkillsAdaptability
Tools & Technologies
CloudWatchDatadogGitHub ActionsDockerTerraformAWSCI/CD PlatformsGitLabJenkinsMonitoring and Alerting Tools
Certifications & Qualifications
AWS DevSecOps Engineer CertificationAWS CertificationsKubernetes Certifications
Industry Keywords
Cloud Security Best PracticesZero Trust Security ModelsHighly Regulated EnvironmentsFederal ProgramsLarge-Scale Enterprise Programs

Tech Stack

Tools & technologies
AWSCloudDockerJavaJenkinsKubernetesPythonTerraformGo

About the role

Key responsibilities & impact
  • Lead infrastructure-level disaster recovery drill execution
  • Validate platform rebuild procedures and disaster recovery playbooks
  • Execute infrastructure-level drills to ensure the platform can be fully rebuilt within the 48-hour recovery target
  • Verify end-to-end data completeness, integrity, and accuracy during drill exercises
  • Document drill results and remediation recommendations
  • Identify exit-readiness gaps across infrastructure, deployment automation, monitoring, and data recovery processes
  • Drive corrective actions with engineering teams
  • Design, implement, and support automated IaC workflows using Terraform, AWS CloudFormation, and standardized CI/CD pipelines
  • Manage and optimize Kubernetes clusters and Docker containerized workloads, including provisioning, scaling, and reliability improvements
  • Build and maintain observability solutions using CloudWatch, Datadog, and other monitoring and alerting tools
  • Develop automation, tooling, and scripts using Python or Java
  • Collaborate with platform engineering, security, applications, and data teams to maintain secure and compliant platform operations
  • Participate in on-call rotations, root cause analyses, and incident response activities

Requirements

What you’ll need
  • Bachelor’s degree and 8–10 years of relevant SRE, DevOps, cloud engineering, or infrastructure engineering experience; or 12 years of experience with a high school diploma
  • Expert-level hands-on knowledge of AWS services across compute, networking, storage, IAM, and serverless components
  • Strong experience with Infrastructure as Code using Terraform and CloudFormation
  • Experience building CI/CD deployment pipelines and progressive delivery mechanisms using GitHub Actions or similar tools
  • Deep understanding of Kubernetes administration, container orchestration, and Docker-based deployments
  • Proven experience validating disaster recovery processes, performing system rebuilds, and conducting data integrity checks
  • Experience building monitoring tools such as dashboards, metrics, logs, and alerting systems using CloudWatch, Datadog, or similar observability tools
  • Proficiency with Python, Java, C#, or Go
  • Experience debugging complex failure modes, including cascading failures, network partitions, backpressure, and eventual consistency issues
  • Strong analytical and documentation skills
  • Ability to work in a fast-paced environment supporting high-visibility, mission-critical systems
  • Ability to obtain a Public Trust clearance
  • US Citizen or Green Card Holder
  • AWS DevSecOps Engineer certification preferred
  • Additional AWS certifications and/or Kubernetes certifications preferred
  • Familiarity with Zero Trust security models and cloud security best practices preferred
  • Experience with GitLab, Jenkins, or similar CI/CD platforms preferred
  • Experience with highly regulated environments preferred
  • Experience supporting federal, defense, or large-scale enterprise programs preferred
  • Prior involvement in large-scale DR drills, COOP, or portability/executable readiness assessments preferred

Benefits

Comp & perks
  • Potential eligibility for overtime
  • Potential eligibility for shift differential
  • Potential eligibility for a discretionary bonus
  • Equal opportunity employment, including disability and protected veterans