FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and developing automation tools, incident management strategies, and data pipelines while leading technical decisions in complex distributed systems. Proficient in multiple programming languages and cloud technologies, with a strong focus on observability, reliability engineering, and incident response.
Highest-signal resume keywords
Go ProgrammingKubernetes Application DevelopmentIncident ManagementObservability EngineeringTechnical Leadership
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Java ProgrammingPython ProgrammingC# ProgrammingSQL TechnologiesNoSQL TechnologiesData Pipeline DevelopmentAnalytics and Dashboard CreationRoot Cause AnalysisSystem DesignSoftware Engineering Fundamentals
Soft Skills
Strong Communication SkillsCoaching and MentoringAbility to Perform Under Pressure
Tools & Technologies
KubernetesAzureAWSGrafanaPower BIDatadogSplunkPagerDutyOpenTelemetryAI-Assisted Development Tools
Industry Keywords
Incident ResponseOperational ReadinessCI/CDInfrastructure as CodePost-Incident ReviewReliability EngineeringCloud-Native ServicesComplex Distributed Systems24x7 Production EnvironmentsTechnical Design Decisions
Tech Stack
Tools & technologiesAWSAzureCloudDistributed SystemsGoogle Cloud PlatformGrafanaJavaKubernetesNoSQLPythonSparkSplunkSQLGo
About the role
Key responsibilities & impact- Design, develop, and operate automation, self-service tools, dashboards, and data pipelines for incident management, on-call, paging, and troubleshooting
- Build shared services, APIs, data contracts, automation, and integrations to standardize incident response and reduce operational risk
- Apply engineering standards across design, implementation, deployment, testing, observability, security, operational support, and production readiness
- Champion safe deployment, CI/CD, infrastructure as code, automated testing, rollback patterns, and operational controls
- Evaluate and implement technologies and tools that improve platform capability, compliance, visibility, reliability, and engineering effectiveness
- Act as a technical leader and escalation point during high-severity incidents
- Guide troubleshooting strategy, cross-team coordination, impact analysis, and risk-based decision making
- Lead or materially contribute to post-incident reviews, root cause analysis, corrective action planning, and systemic reliability improvements
- Develop and maintain incident response strategies, operational runbooks, readiness criteria, triage models, and resilience practices
- Lead complex design and architecture reviews across teams, services, dependencies, and operational domains
- Partner with SRE, platform, product, infrastructure, security, and business stakeholders
- Translate technical concepts, risks, and tradeoffs for technical leaders, business stakeholders, and non-technical audiences
- Mentor senior and mid-level engineers through technical leadership, code and design reviews, documentation, and operational coaching
- Participate in a 24x7 on-call rotation supporting incident response and production support
Requirements
What you’ll need- Hands-on proficiency in multiple languages, including Go, Java, Python, and C#
- Experience building production-grade full-stack applications on Kubernetes and serverless technologies such as Knative in Azure and AWS
- Experience with SQL and NoSQL technologies and cloud-native services
- Experience building and using data pipelines, analytics, and dashboards using Spark, Trino, Grafana, Superset, and Power BI
- Experience with OpenTelemetry and observability platforms such as Grafana, Datadog, Splunk, and Azure Monitor
- Experience with incident management platforms such as PagerDuty
- Proficiency with AI-assisted development tools such as Claude Code, Cursor, and GitHub Copilot
- Experience improving incident, post-incident review, or reliability processes
- Strong incident forensics and root cause analysis skills
- Strong understanding of observability, reliability engineering, incident management, and post-incident improvement practices
- Experience supporting high-severity production incidents in complex environments
- Strong software engineering fundamentals and system design skills
- Ability to lead technical design decisions in complex distributed systems
- Strong communication skills and ability to coach engineers and present findings clearly to leadership
- 8+ years of professional software engineering experience
- 6+ years of experience with architecture, design, system reliability, scalability, and technical leadership for production systems
- 4+ years of experience with open-source frameworks, modern engineering practices, or platform technologies
- 3+ years of experience with Azure, AWS, GCP, or another cloud service provider, or equivalent experience in complex hybrid environments
- Demonstrated ownership of mission-critical systems operating in 24x7 production environments
- Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience
- Ability to influence engineering outcomes across teams
- Must communicate in a clear, concise, professional oral and written manner
- Must perform effectively under pressure and in stressful situations
- Must participate in a 24x7 on-call rotation
- At this time, GEICO will not sponsor a new applicant for employment authorization for this position
Benefits
Comp & perks- Personalized development programs
- Mentorship
- Certification assistance
- Competitive pay
- Benefits
- Flexibility to support well-being and future
- Equal employment opportunity
- Disability accommodations
