FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in site reliability engineering, DevOps practices, and infrastructure management, with a strong focus on AWS and Azure environments. Proficient in defining SLOs, incident response, and automation to enhance system reliability and compliance.
Highest-signal resume keywords
Site Reliability EngineeringAWS Services ManagementAzure Infrastructure ManagementInfrastructure-as-Code with TerraformIncident Response Leadership
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
SQL Database AdministrationDocker DeploymentPowerShell ScriptingCI/CD PracticesMonitoring and AlertingAutomation of Repetitive TasksError Budget ManagementRoot Cause AnalysisCompliance RequirementsService Level Objectives
Soft Skills
Mentoring EngineersCross-Team CollaborationProblem-Solving
Tools & Technologies
AWSAzureTerraformDockerWindows VMsIISEventBridgeSNSSQSRDS
Certifications & Qualifications
AZ-104AZ-305AWS Certified Solutions Architect
Industry Keywords
HIPAA ComplianceFedRAMPPublic TrustDoDVA Health ProgramsZero-Trust ArchitectureITIL Service Management12-Factor App
Tech Stack
Tools & technologiesAWSAzureCloudDockerSQLTerraform
About the role
Key responsibilities & impact- Own the reliability of production systems and how reliability is measured
- Define SLOs and SLIs with stakeholders, establish error budgets, and use them to guide release decisions
- Run and improve the SaaS monitoring, alerting, and logging stack to meet compliance requirements and detect issues proactively
- Participate in on-call production incident response and help engineers resolve customer issues
- Lead blameless postmortems, identify root causes, and implement lasting fixes
- Identify, measure, and automate manual repetitive engineering work
- Maintain automation and documentation that survives ownership changes
- Own the design, build, and maintenance of core infrastructure for safe and efficient product delivery
- Set priorities for reliability and platform work, obtain manager approval, and adapt to business needs
- Assess risk versus impact for high-visibility systems and roll out measurable, reversible improvements
- Make build-versus-buy decisions and recommend proven tools
- Partner with engineering teams on CI/CD, testing, canary releases, and automated rollback practices
- Work with security to maintain infrastructure and operations compliance and proactively raise risks
- Mentor engineers on reliability, Azure, AWS, Terraform, and operational practices
- Create documentation for system operation and lead solutions with clear tradeoffs
Requirements
What you’ll need- 8+ years in site reliability, DevOps, or infrastructure engineering, including ownership of production systems from end to end
- Built and managed AWS and Azure accounts and resources in line with the Well-Architected Framework
- Experience with AWS services including ECS, Fargate, RDS, Lambda, SNS, SQS, S3, EventBridge, and Step Functions
- Deployed Docker-based software to production and understanding of reliable container operations
- Hands-on experience across Azure, AWS, and traditional data centers, including managing Windows VMs and IIS
- Deep expertise in SQL and relational database administration, including query tuning, index optimization, and resolving high-load production incidents; SQL Server preferred
- Infrastructure-as-code experience using Terraform
- Practical experience with DevOps principles, the 12-Factor App, least privilege access, and zero-trust architecture
- Experience leading cross-team requirements and SLA definition
- Experience defining and operating against SLOs and leading incident response and postmortems for production outages
- Experience with ITIL-aligned service management, including incident, problem, and change management
- Experience mentoring engineers on reliability and delivery practices
- United States Citizenship
- Ability to satisfy security investigation and eligibility requirements for access to classified (Public Trust) information
- Azure or AWS certifications preferred, such as AZ-104, AZ-305, or AWS Certified Solutions Architect
- Experience with Azure Government or other federal cloud environments preferred
- Experience supporting the VA, DoD, or other federal health programs, including FedRAMP or ATO work preferred
- PowerShell scripting and automation preferred
- Experience in a HIPAA-compliant environment preferred
- Security experience beyond day-to-day operations, such as red teaming or penetration testing, preferred
Benefits
Comp & perks- Flexible work schedule
- Unlimited PTO
- Physical and mental wellness benefits
- Medical coverage
- Parental leave
- 401K
- Company-sponsored events
- Referral program
- Onsite gym
- Dog friendly office
- Snacks in the office
- Commuter benefits
- Onsite massages
