FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Staff Engineer – SRE, Incident Prevention, Post Incident Correction of Errors
GEICO. Design, develop, and operate automation, self-service tools, dashboards, and data pipelines that automate and scale COE workflows .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in automation, incident management, and system design, with a strong focus on improving reliability and scalability in production environments. Proficient in multiple programming languages and cloud technologies, with a proven ability to lead technical initiatives and coach engineering teams.
Highest-signal resume keywords
Hands-On Proficiency In Go, Java, Python, And C#Experience With Kubernetes And Serverless TechnologiesDeep Incident Forensics And Root Cause Analysis SkillsStrong Understanding Of Observability And Reliability Engineering10+ Years Of Professional Software Engineering Experience
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AutomationIncident ManagementSystem DesignRoot Cause AnalysisData PipelinesFull-Stack Application DevelopmentObservabilityTechnical LeadershipSoftware Engineering FundamentalsCloud Technologies
Soft Skills
Strong Communication SkillsCoaching Engineering TeamsAbility To Perform Under PressureInfluencing Engineering Outcomes
Tools & Technologies
KubernetesAzureAWSGrafanaDatadogSplunkPower BISQLNoSQLPagerDuty
Industry Keywords
Incident ResponseHigh-Severity IncidentsProduction SupportContinuous ImprovementMission-Critical Systems
Tech Stack
Tools & technologiesAWSAzureCloudDistributed SystemsGoogle Cloud PlatformGrafanaJavaKubernetesNoSQLPythonSparkSplunkSQLGo
About the role
Key responsibilities & impact- Design, develop, and operate automation, self-service tools, dashboards, and data pipelines that automate and scale COE workflows
- Reduce manual tracking and follow-ups
- Run and moderate weekly GEICO-wide COE presentation sessions for qualified high-severity incidents
- Lead and improve the COE process across the engineering organization
- Provide technical leadership in system design, architecture, hands-on engineering, COE improvements, automation, and incident tooling
- Coach application engineering teams on identifying true root causes and producing clear, complete, high-quality COEs
- Lead root cause analysis across distributed systems using logs, metrics, traces, and observability data
- Identify gaps such as missing alerts, weak monitoring, incomplete runbooks, poor testing, and bypassed deployment controls
- Analyze incidents for repeat patterns and connect lessons across COEs
- Share insights that improve prevention and reliability
- Partner with application engineering, platform, SRE, and operations stakeholders to drive accountability and continuous improvement
- Improve technology strategies and roadmaps
- Participate in high-severity incident leadership and management
- Participate in a 24x7 on-call rotation supporting incident response and production support for mission-critical platforms
Requirements
What you’ll need- Hands-on proficiency in multiple languages, including Go, Java, Python, and C#
- Experience building production-grade full-stack applications on Kubernetes and serverless technologies (KNative) in Azure and AWS
- Experience with SQL and NoSQL technologies
- Experience building and using data pipelines, analytics, and dashboards using Spark, Trino, Grafana, Superset, and Power BI
- Experience with OpenTelemetry and observability platforms such as Grafana, Datadog, Splunk, and Azure Monitor
- Experience with incident management platforms such as PagerDuty
- Proficiency with AI-assisted development tools such as Claude Code, Cursor, and GitHub Copilot
- Experience improving incident, post-incident review, or reliability processes at scale through automation, data, and cross-team influence
- Deep incident forensics and root cause analysis skills
- Strong understanding of observability, reliability engineering, incident management, and post-incident improvement practices
- Experience supporting incident response and high-severity production incidents in complex environments
- Strong software engineering fundamentals and system design skills
- Ability to lead technical design and architecture decisions in complex distributed systems
- Strong communication skills and ability to coach engineering teams and present findings clearly to leadership
- 10+ years of professional software engineering experience
- 8+ years of experience with architecture, design, system reliability, scalability, and technical leadership for production systems
- 6+ years of experience with open-source frameworks, modern engineering practices, or platform technologies
- 4+ years of experience with Azure, AWS, GCP, or another cloud service provider, or equivalent experience in complex hybrid environments
- Demonstrated ownership of mission-critical systems operating in 24x7 production environments
- Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience
- Ability to influence engineering outcomes across teams
- Clear, concise, professional oral and written communication
- Ability to perform effectively under pressure, including during production support and high-severity incident response
- Ability to participate in a 24x7 on-call rotation
- GEICO will not sponsor a new applicant for employment authorization for this position
Benefits
Comp & perks- Personalized development programs
- Mentorship
- Certification assistance
- Competitive pay
- Benefits
- Flexibility to support well-being and future
- Inclusive and collaborative culture