FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Reliability Engineer
SES Corporation. Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments .
Posted 9/26/2026full-timeHanscom Air Force Base • Massachusetts • United StatesMid-LevelSeniorWebsite
Tech Stack
Tools & technologiesAWSAzureCloudDistributed SystemsDNSDockerGoogle Cloud PlatformGrafanaKubernetesLinuxOraclePrometheusPythonTCP/IPTerraformGo
About the role
Key responsibilities & impact- Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments
- Define, measure, and report SLIs, SLOs, and error budgets
- Identify reliability risks and implement mitigation strategies across the system lifecycle
- Conduct capacity planning and performance modeling
- Implement and manage monitoring, logging, and tracing solutions
- Define actionable alerting thresholds and analyze reliability trends and metrics
- Participate in on-call rotations and lead production incident response
- Coordinate troubleshooting across development, infrastructure, and security teams
- Conduct post-incident reviews and develop corrective and preventive action plans
- Track recurring issues and ensure root causes are resolved
- Automate operational tasks and develop scripts, tools, and services to improve reliability and reduce MTTR
- Participate in architecture and design reviews focused on reliability, resiliency, and recoverability
- Validate disaster recovery and business continuity plans; test failover mechanisms
- Support chaos engineering, fault injection testing, and resilience validation
- Partner with DevOps, Platform, and Security teams
- Document reliability standards, runbooks, and operational procedures
- Support compliance and audit activities including FedRAMP, FISMA, and internal operational controls
Requirements
What you’ll need- Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience; additional experience may be accepted in lieu of degree
- Active Secret clearance at a minimum required to start
- US citizenship required
- Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services
- Experience with containerized environments (Docker, Kubernetes)
- Familiarity with CI/CD pipelines and deployment automation
- Experience with SLOs and error budgets
- Experience with capacity modeling and performance testing
- Strong understanding of distributed systems and high-availability architectures
- Strong understanding of Linux/Windows system administration
- Strong understanding of networking fundamentals (DNS, TCP/IP, load balancing)
- Hands-on experience with monitoring and observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor)
- Hands-on experience with Infrastructure as Code (Terraform, ARM, CloudFormation)
- Hands-on experience with scripting or programming languages (Python, Bash, Go, PowerShell, or similar)
- Experience supporting incident management and on-call operations
- Preferred: Experience with USAF Cloud One or Platform 1
- Preferred: Experience with Zero Trust Architecture
- Preferred: Cloud certifications in AWS, Azure, Google, or Oracle clouds
Benefits
Comp & perks- Medical
- Dental
- Vision
- AD&D
- STD
- LTD
- Company paid Life Insurance
- 401k with employer contribution
- Paid Time Off
- Pet Insurance