Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Coalfire

Senior Site Reliability Engineer – Fedramp

Coalfire

. Own an operational capability for the managed estate, including automation, runbooks, and service standards .

Posted 9/17/2026full-timeRemote • United StatesSenior💰 $85,000 - $141,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering and cloud operations, with a strong focus on automation, observability, and incident management. Proven ability to mentor teams, lead technical discussions, and implement effective operational capabilities in regulated environments.

Highest-signal resume keywords
Site Reliability EngineeringInfrastructure-as-CodeCloud Operations (AWS, Azure, GCP)Incident Response and CommandObservability Engineering

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
AutomationBackup and Recovery EngineeringContinuous MonitoringTelemetry and Log PipelinesService-Level ObjectivesIncident ManagementScriptingPolicy-as-CodeTechnical DocumentationOperational Capability Ownership
Soft Skills
Excellent CommunicationOrganizational SkillsProblem-Solving SkillsCritical ThinkingMentoring
Tools & Technologies
TerraformAnsibleCI/CD PipelinesMonitoring PlatformsRunbooks
Certifications & Qualifications
AWS CertificationAzure CertificationGCP Certification
Industry Keywords
NIST 800-53FedRAMPOperational PostureIncident CommandRegulated Cloud Environments

Tech Stack

Tools & technologies
AnsibleAWSAzureCloudGoogle Cloud PlatformTerraform

About the role

Key responsibilities & impact
  • Own an operational capability for the managed estate, including automation, runbooks, and service standards
  • Design observability for regulated cloud environments, including telemetry and log pipelines, service-level objectives, alert quality, and escalation paths
  • Build and maintain continuous-monitoring evidence pipelines
  • Own backup and recovery engineering, including tested recovery procedures, measurable recovery objectives, and outage automation
  • Serve as the senior escalation point in client environments and resolve complex operational incidents
  • Lead incident and problem management, including major-event incident command, blameless post-incident reviews, and corrective actions
  • Automate operational toil using infrastructure-as-code, pipelines, and scripting
  • Partner with Engagement Architects and Build teams on transition into managed operations
  • Hold on-call responsibility and improve rotation coverage, alert actionability, and team load
  • Represent operational posture to clients and support renewals and expansions
  • Mentor and lead Site Reliability Engineers and junior staff
  • Author and peer review code, runbooks, operational design documentation, and compliance artifacts

Requirements

What you’ll need
  • BS or above in a related Information Technology field or equivalent combination of education and experience
  • Bachelor’s degree or equivalent combination of education and work experience
  • Professional- or specialty-level certification in AWS, Azure, or GCP; associate-level certification considered with equivalent demonstrated depth
  • 5+ years in site reliability engineering, cloud operations, platform engineering, or managed services
  • 5+ years operating production cloud environments in AWS, Azure, or GCP, including monitoring, incident response, and automation
  • Automation-first mindset with deep Infrastructure-as-Code, CI/CD, scripting, and policy-as-code
  • Deep operational command of at least one major cloud platform and working knowledge of a second
  • Observability engineering experience with metrics, logging, log pipelines, distributed tracing, SLI/SLO definition, and alert design
  • Demonstrated incident response and incident command capability
  • Backup, recovery, and resilience engineering experience
  • Working command of NIST 800-53, FedRAMP, or comparable security control frameworks
  • Ability to lead technical client conversations about operational posture, risk, and trade-offs
  • Demonstrated ability to mentor engineers and improve team output
  • Excellent communication, organizational, and problem-solving skills
  • Effective documentation skills, including technical diagrams, runbooks, and written descriptions
  • Ability to work independently and as part of a team
  • Critical thinking and ability to balance security and availability requirements against mission needs
  • Demonstrated experience owning an operational capability, monitoring platform, or reusable automation used by multiple teams or clients
  • Experience as the senior operational escalation point on client-facing managed services, including incident command on major events
  • Advanced experience with Infrastructure-as-Code and orchestration/automation tools such as Terraform and Ansible
  • Experience transitioning environments from build into steady-state operations

Benefits

Comp & perks
  • Flexible work model allowing employees to choose when and where they work
  • Paid parental leave
  • Flexible time off
  • Certification and training reimbursement
  • Digital mental health and wellbeing support membership
  • Comprehensive insurance options
  • Employee resource groups
  • In-person and virtual events
  • Annual incentive, commission, and/or recognition programs may be available