FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering and cloud operations, with a strong focus on automation, observability, and incident management. Proven ability to mentor teams, lead technical discussions, and implement effective operational capabilities in regulated environments.
Highest-signal resume keywords
Site Reliability EngineeringInfrastructure-as-CodeCloud Operations (AWS, Azure, GCP)Incident Response and CommandObservability Engineering
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AutomationBackup and Recovery EngineeringContinuous MonitoringTelemetry and Log PipelinesService-Level ObjectivesIncident ManagementScriptingPolicy-as-CodeTechnical DocumentationOperational Capability Ownership
Soft Skills
Excellent CommunicationOrganizational SkillsProblem-Solving SkillsCritical ThinkingMentoring
Tools & Technologies
TerraformAnsibleCI/CD PipelinesMonitoring PlatformsRunbooks
Certifications & Qualifications
AWS CertificationAzure CertificationGCP Certification
Industry Keywords
NIST 800-53FedRAMPOperational PostureIncident CommandRegulated Cloud Environments
Tech Stack
Tools & technologiesAnsibleAWSAzureCloudGoogle Cloud PlatformTerraform
About the role
Key responsibilities & impact- Own an operational capability for the managed estate, including automation, runbooks, and service standards
- Design observability for regulated cloud environments, including telemetry and log pipelines, service-level objectives, alert quality, and escalation paths
- Build and maintain continuous-monitoring evidence pipelines
- Own backup and recovery engineering, including tested recovery procedures, measurable recovery objectives, and outage automation
- Serve as the senior escalation point in client environments and resolve complex operational incidents
- Lead incident and problem management, including major-event incident command, blameless post-incident reviews, and corrective actions
- Automate operational toil using infrastructure-as-code, pipelines, and scripting
- Partner with Engagement Architects and Build teams on transition into managed operations
- Hold on-call responsibility and improve rotation coverage, alert actionability, and team load
- Represent operational posture to clients and support renewals and expansions
- Mentor and lead Site Reliability Engineers and junior staff
- Author and peer review code, runbooks, operational design documentation, and compliance artifacts
Requirements
What you’ll need- BS or above in a related Information Technology field or equivalent combination of education and experience
- Bachelor’s degree or equivalent combination of education and work experience
- Professional- or specialty-level certification in AWS, Azure, or GCP; associate-level certification considered with equivalent demonstrated depth
- 5+ years in site reliability engineering, cloud operations, platform engineering, or managed services
- 5+ years operating production cloud environments in AWS, Azure, or GCP, including monitoring, incident response, and automation
- Automation-first mindset with deep Infrastructure-as-Code, CI/CD, scripting, and policy-as-code
- Deep operational command of at least one major cloud platform and working knowledge of a second
- Observability engineering experience with metrics, logging, log pipelines, distributed tracing, SLI/SLO definition, and alert design
- Demonstrated incident response and incident command capability
- Backup, recovery, and resilience engineering experience
- Working command of NIST 800-53, FedRAMP, or comparable security control frameworks
- Ability to lead technical client conversations about operational posture, risk, and trade-offs
- Demonstrated ability to mentor engineers and improve team output
- Excellent communication, organizational, and problem-solving skills
- Effective documentation skills, including technical diagrams, runbooks, and written descriptions
- Ability to work independently and as part of a team
- Critical thinking and ability to balance security and availability requirements against mission needs
- Demonstrated experience owning an operational capability, monitoring platform, or reusable automation used by multiple teams or clients
- Experience as the senior operational escalation point on client-facing managed services, including incident command on major events
- Advanced experience with Infrastructure-as-Code and orchestration/automation tools such as Terraform and Ansible
- Experience transitioning environments from build into steady-state operations
Benefits
Comp & perks- Flexible work model allowing employees to choose when and where they work
- Paid parental leave
- Flexible time off
- Certification and training reimbursement
- Digital mental health and wellbeing support membership
- Comprehensive insurance options
- Employee resource groups
- In-person and virtual events
- Annual incentive, commission, and/or recognition programs may be available
