Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Peraton

Site Reliability Engineer – Night Shift

Peraton

. Operate and maintain production infrastructure services and applications, ensuring availability, reliability, performance, security, and operational health .

Posted 9/18/2026full-timeRemote • United StatesSeniorLead💰 $104,000 - $166,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in site reliability engineering and DevOps practices, with a strong focus on operational health, incident management, and automation. Proficient in managing production infrastructure and implementing observability metrics to enhance service reliability and performance.

Highest-signal resume keywords
Site Reliability EngineeringInfrastructure-As-CodeAWS Commercial and GovCloudCI/CD PlatformsEnterprise Observability Tools

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Site Reliability EngineeringInfrastructure-As-CodeAWS CommercialAWS GovCloudTerraformAnsibleCI/CDLinux AdministrationWindows Server AdministrationScripting
Soft Skills
CollaborationProblem-SolvingOwnership
Tools & Technologies
GitLabJenkinsDynatraceDatadogSplunkOpen TelemetryOpenShiftKubernetes
Certifications & Qualifications
AWS Solutions ArchitectAWS DevOps EngineerAWS SysOpsRed Hat Certified Specialist in ROSARed Hat Certified System AdministratorAzure Administrator AssociateGCP Associate Cloud EngineerDynatrace AssociateDatadog Log Management FundamentalsGitLab CI/CD Associate
Industry Keywords
FISMAFedRAMPNIST 800-53Operational HealthIncident ManagementService ReliabilityCapacity PlanningDisaster RecoveryPerformance TestingOperational Technical Debt

Tech Stack

Tools & technologies
AnsibleAWSAzureCloudGoogle Cloud PlatformJenkinsKubernetesLinuxOpenShiftPythonSplunkTerraformGo

About the role

Key responsibilities & impact
  • Operate and maintain production infrastructure services and applications, ensuring availability, reliability, performance, security, and operational health
  • Monitor services and applications using SLIs, SLOs, dashboards, alerts, and observability tools
  • Improve detection, diagnosis, and resolution of operational issues
  • Partner with application teams to define observability requirements and implement metrics, logs, traces, dashboards, and alerts
  • Manage production incidents and service disruptions, including on-call response, troubleshooting, service restoration, root-cause analysis, and post-incident corrective actions
  • Execute application and infrastructure releases through deployment pipelines, including staging and production promotion, validation, rollback, and release troubleshooting
  • Manage the operational lifecycle of infrastructure, including upgrades, patching, configuration changes, maintenance, and technology refreshes
  • Assess and improve service resilience through capacity planning, performance testing, failure-mode analysis, disaster recovery, backup, failover, and recovery testing
  • Identify and address reliability risks and operational technical debt using reliability metrics, incident trends, capacity data, and service health indicators
  • Automate operational activities using an everything-as-code approach
  • Collaborate with platform engineering and application teams to identify operational requirements and improve environment reliability and operability

Requirements

What you’ll need
  • Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level clearance
  • Bachelor's Degree and 8 years of experience, or a High School diploma/equivalent and 12 years of experience
  • 7+ years hands-on experience in site reliability engineering, DevOps, or production systems engineering
  • Hands-on experience operating in AWS Commercial and AWS GovCloud, including OpenShift (ROSA) or comparable Kubernetes-based platforms
  • Strong infrastructure-as-code experience with Terraform and Ansible/Ansible Tower
  • Experience with CI/CD platforms GitLab and Jenkins, including reliability gating and deployment automation
  • Proficient in Linux and Windows Server administration
  • Experience with enterprise observability tools such as Dynatrace, Datadog, Splunk and Open Telemetry
  • Demonstrated ownership of an SLI/SLO and alerting program, including error budgets, alert rationalization, and noise reduction
  • Scripting/automation proficiency in Python, Bash, PowerShell, or Go
  • Experience operating in federal or regulated environments (FISMA, FedRAMP, NIST 800-53)
  • Preferred: AWS Solutions Architect, AWS DevOps Engineer, or AWS SysOps certification
  • Preferred: Red Hat Certified Specialist in ROSA or Red Hat Certified System Administrator in OpenShift
  • Preferred: Azure Administrator Associate or GCP Associate Cloud Engineer certification
  • Preferred: Dynatrace Associate or Datadog Log Management Fundamentals certification
  • Preferred: GitLab CI/CD Associate certification or Certified Jenkins Engineer (CJE)
  • Preferred: Terraform Associate certification

Benefits

Comp & perks
  • Eligible for overtime
  • Shift differential may be available
  • Discretionary bonus may be available