FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive experience in Site Reliability Engineering, focusing on production web stack operability, incident response, and configuration management. Proven ability to lead teams, drive operational excellence, and implement monitoring and alerting systems for high-availability environments.
Highest-signal resume keywords
Site Reliability EngineeringProduction Web Stack ManagementIncident Response LeadershipConfiguration ManagementMonitoring and Alerting Systems
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Linux Systems KnowledgeDatabase ReplicationPerformance TuningDebugging Performance IssuesInfrastructure TransitionsCapacity PlanningDisaster RecoveryObservabilityRunbook DevelopmentSLO Definition and Tracking
Soft Skills
Team LeadershipMentoringWritten CommunicationVerbal CommunicationCross-Functional Coordination
Tools & Technologies
PuppetAnsibleChefRedisHAProxyCloudflareHarvesterKeepalivedLoad Balancing StrategiesMonitoring Tools
Industry Keywords
Production Control PlaneOperational ExcellenceBlameless PostmortemsPCI ComplianceCardholder Data
Tech Stack
Tools & technologiesAnsibleChefDNSHAProxyLinuxPHPPuppetRedis
About the role
Key responsibilities & impact- Build and lead the Production Site Reliability Engineering team
- Hire, mentor, and grow SREs and database engineers responsible for the production control plane
- Own availability, performance, and operability of the production web stack
- Lead incident response for production-impacting events
- Drive postmortems and convert lessons into system improvements, runbooks, and alerting
- Own configuration management for the production environment
- Partner with engineering teams to ensure new services are observable, deployable, and documented before reaching customers
- Drive control plane re-architecture, including new infrastructure, parallel runs, cutover, and production-readiness validation
- Set the operational roadmap covering capacity planning, scaling, disaster recovery, and architectural evolution
- Establish and maintain monitoring, alerting, and observability across the production stack
- Manage production database and caching environments, including replication topology, performance tuning, backup verification, and failover testing
- Foster operational excellence through operational reviews, SLO definition and tracking, blameless postmortems, and continuous improvement of runbooks and deployment processes
Requirements
What you’ll need- 10+ years of professional experience in site reliability engineering or infrastructure operations, with at least 2 years in a team lead or management role
- Deep experience operating production web stacks at scale
- Comfortable debugging performance issues across the full request path from load balancer to database
- Strong Linux systems knowledge: networking, systemd, package management, firewall configuration, and performance tuning
- Track record of building and operating monitoring and alerting systems and driving observability and proactive incident response
- Experience with configuration management at scale, such as Puppet, Ansible, Chef, or similar
- Strong written and verbal communication for runbooks, postmortems, incident communications, and cross-functional coordination
- Experience hiring and building engineering teams
- Experience with database replication, backup strategies, and failure modes
- Experience leading production migrations or major infrastructure transitions
- Bonus: experience with Harvester or similar HCI platforms for running production workloads on VMs
- Bonus: experience with Redis cluster architecture, failover, and performance tuning at scale
- Bonus: experience with HAProxy configuration, keepalived, and load balancing strategies
- Bonus: experience with PCI compliance environments and systems handling cardholder data
- Bonus: experience writing PHP
- Bonus: experience with Cloudflare or similar CDN/edge platforms and DNS cutover
- No specific educational credential stated
Benefits
Comp & perks- 100% company-paid insurance premiums for employee medical, dental and vision plans
- 401(k) plan that matches 100% up to 4%, with immediate vesting
- Professional Development Reimbursement of $2,500 each year
- 11 Holidays + Paid Time Off Accrual + Rollover Plan
- Increased PTO at 3 year and 10 year anniversary
- 1 month paid sabbatical every 5 years
- Anniversary Bonus each year
- $500 stipend for remote office setup in first year + $400 each following year
- Internet reimbursement up to $75 per month
- Gym membership reimbursement up to $50 per month
- Company paid Wellable subscription
