FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in incident management, observability, and reliability engineering, with a strong focus on building and maintaining automation tools, data pipelines, and production systems. Proven ability to lead technical design decisions and mentor engineering teams while effectively communicating complex concepts to diverse stakeholders.
Highest-signal resume keywords
Go ProgrammingKubernetes Application DevelopmentIncident ManagementObservability EngineeringCloud Services (Azure, AWS)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Java ProgrammingPython ProgrammingC# ProgrammingSQL TechnologiesNoSQL TechnologiesData Pipeline DevelopmentRoot Cause AnalysisSystem DesignCI/CD PracticesAI-Assisted Development Tools
Soft Skills
Strong Communication SkillsCoaching Engineering TeamsAbility to Perform Under Pressure
Tools & Technologies
GrafanaDatadogSplunkAzure MonitorPagerDutySparkTrinoPowerBIKNativeOpenTelemetry
Industry Keywords
Incident ResponseOperational ReadinessProduction SystemsSystem ReliabilityPost-Incident Review
Tech Stack
Tools & technologiesAWSAzureCloudDistributed SystemsGoogle Cloud PlatformGrafanaJavaKubernetesNoSQLPythonSparkSplunkSQLGo
About the role
Key responsibilities & impact- Design, develop, and operate automation, self-service tools, dashboards, and data pipelines for incident management, on-call, paging, and troubleshooting processes
- Build shared services, APIs, data contracts, automation, and integrations to standardize incident response and reduce operational risk
- Set and uphold engineering standards across design, implementation, deployment, testing, observability, security, operational support, and production readiness
- Champion safe deployment, CI/CD, infrastructure as code, automated testing, rollback patterns, and operational controls
- Evaluate, select, and implement modern technologies and tools to improve platform capability, compliance, visibility, reliability, and engineering effectiveness
- Act as a technical leader and escalation point during high-severity incidents
- Guide troubleshooting strategy, cross-team coordination, impact analysis, and risk-based decision-making
- Lead or heavily influence post-incident reviews, root cause analysis, corrective action planning, and systemic reliability improvements
- Develop and maintain incident response strategies, operational runbooks, readiness criteria, triage models, and resilience practices
- Lead complex design and architecture reviews across multiple teams, services, dependencies, and operational domains
- Partner with SRE, platform, product, infrastructure, security, and business stakeholders
- Translate technical concepts, risks, and tradeoffs for executives, technical leaders, and non-technical stakeholders
- Mentor senior engineers, staff engineers, and teams through technical leadership, code/design reviews, documentation, and operational coaching
- Participate in a 24x7 on-call rotation supporting incident response and production support
Requirements
What you’ll need- Hands-on proficiency in multiple languages, including Go, Java, Python, and C#
- Experience building production-grade full-stack applications on Kubernetes and serverless technologies (KNative) in Azure and AWS
- Experience with SQL and NoSQL technologies and cloud-native services
- Experience building and using data pipelines, analytics, and dashboards using Spark, Trino, Grafana, Superset, and PowerBI
- Experience with OpenTelemetry and observability platforms such as Grafana, Datadog, Splunk, and Azure Monitor
- Experience with incident management platforms such as PagerDuty
- Proficiency with AI-assisted development tools such as Claude Code, Cursor, and GitHub Copilot
- Experience improving incident, post-incident review, or reliability processes at scale
- Deep incident forensics and root cause analysis skills
- Strong understanding of observability, reliability engineering, incident management, and post-incident improvement practices
- Experience supporting high-severity production incidents in complex environments
- Strong software engineering fundamentals and system design skills
- Ability to lead technical design and architecture decisions in complex distributed systems
- Strong communication skills and ability to coach engineering teams and present findings clearly to leadership
- 10+ years of professional software engineering experience
- 8+ years of experience with architecture, design, system reliability, scalability, and technical leadership for production systems
- 6+ years of experience with open-source frameworks, modern engineering practices, or platform technologies
- 4+ years of experience with Azure, AWS, GCP, or another cloud service provider, or equivalent experience in complex hybrid environments
- Demonstrated ownership of mission-critical systems operating in 24x7 production environments
- Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience
- Ability to influence engineering outcomes across teams
- Clear, concise, professional oral and written communication
- Ability to perform effectively under pressure and in stressful situations
- Ability to participate in a 24x7 on-call rotation
- GEICO will not sponsor a new applicant for employment authorization for this position
Benefits
Comp & perks- Personalized development programs
- Mentorship
- Certification assistance
- Inclusive and collaborative culture
- Competitive pay
- Benefits
- Flexibility to support well-being and future
- Reasonable accommodations for qualified individuals with disabilities
