Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
SES Corporation

Reliability Engineer

SES Corporation

. Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments .

Posted 9/26/2026full-timeHanscom Air Force Base • Massachusetts • United StatesMid-LevelSeniorWebsite

Tech Stack

Tools & technologies
AWSAzureCloudDistributed SystemsDNSDockerGoogle Cloud PlatformGrafanaKubernetesLinuxOraclePrometheusPythonTCP/IPTerraformGo

About the role

Key responsibilities & impact
  • Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments
  • Define, measure, and report SLIs, SLOs, and error budgets
  • Identify reliability risks and implement mitigation strategies across the system lifecycle
  • Conduct capacity planning and performance modeling
  • Implement and manage monitoring, logging, and tracing solutions
  • Define actionable alerting thresholds and analyze reliability trends and metrics
  • Participate in on-call rotations and lead production incident response
  • Coordinate troubleshooting across development, infrastructure, and security teams
  • Conduct post-incident reviews and develop corrective and preventive action plans
  • Track recurring issues and ensure root causes are resolved
  • Automate operational tasks and develop scripts, tools, and services to improve reliability and reduce MTTR
  • Participate in architecture and design reviews focused on reliability, resiliency, and recoverability
  • Validate disaster recovery and business continuity plans; test failover mechanisms
  • Support chaos engineering, fault injection testing, and resilience validation
  • Partner with DevOps, Platform, and Security teams
  • Document reliability standards, runbooks, and operational procedures
  • Support compliance and audit activities including FedRAMP, FISMA, and internal operational controls

Requirements

What you’ll need
  • Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience; additional experience may be accepted in lieu of degree
  • Active Secret clearance at a minimum required to start
  • US citizenship required
  • Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services
  • Experience with containerized environments (Docker, Kubernetes)
  • Familiarity with CI/CD pipelines and deployment automation
  • Experience with SLOs and error budgets
  • Experience with capacity modeling and performance testing
  • Strong understanding of distributed systems and high-availability architectures
  • Strong understanding of Linux/Windows system administration
  • Strong understanding of networking fundamentals (DNS, TCP/IP, load balancing)
  • Hands-on experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor)
  • Hands-on experience with Infrastructure as Code (Terraform, ARM, CloudFormation)
  • Hands-on experience with scripting or programming languages (Python, Bash, Go, PowerShell, or similar)
  • Experience supporting incident management and on-call operations
  • Preferred: Experience with USAF Cloud One or Platform 1
  • Preferred: Experience with Zero Trust Architecture
  • Preferred: Cloud certifications in AWS, Azure, Google, or Oracle clouds

Benefits

Comp & perks
  • Medical
  • Dental
  • Vision
  • AD&D
  • STD
  • LTD
  • Company paid Life Insurance
  • 401k with employer contribution
  • Paid Time Off
  • Pet Insurance