FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating automation tools, data pipelines, and incident management systems while ensuring compliance and reliability in production environments. Proficient in technical leadership, incident response, and cross-team collaboration to enhance operational effectiveness.
Highest-signal resume keywords
Hands-On Proficiency In Go, Java, Python, Or C#Experience With Kubernetes And Serverless TechnologiesStrong Incident Forensics And Root Cause Analysis SkillsExperience With OpenTelemetry And Observability Platforms4+ Years Of Professional Software Engineering Experience
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AutomationData PipelinesIncident ManagementCI/CDInfrastructure As CodeSQLNoSQLSoftware Engineering FundamentalsSystem DesignTechnical Leadership
Soft Skills
Strong Communication SkillsAbility To Perform Effectively Under PressureMentoring EngineersCross-Team CollaborationClear And Concise Presentation Skills
Tools & Technologies
KubernetesAzureAWSGrafanaDatadogSplunkPower BIPagerDutyClaude CodeGitHub Copilot
Industry Keywords
Incident ResponsePost-Incident ReviewReliability EngineeringObservabilityProduction Systems
Tech Stack
Tools & technologiesAWSAzureCloudDistributed SystemsGoogle Cloud PlatformGrafanaJavaKubernetesNoSQLPythonSparkSplunkSQLGo
About the role
Key responsibilities & impact- Design, develop, and operate automation, self-service tools, dashboards, and data pipelines for incident management, on-call, paging, and troubleshooting
- Build shared services, APIs, data contracts, automation, and integrations to standardize incident response and reduce operational risk
- Apply engineering standards across design, implementation, deployment, testing, observability, security, operational support, and production readiness
- Implement safe deployment, CI/CD, infrastructure as code, automated testing, rollback patterns, and operational controls
- Evaluate and implement technologies and tools that improve platform capability, compliance, visibility, reliability, and engineering effectiveness
- Act as a technical leader during high-severity incidents
- Contribute to troubleshooting strategy, cross-team coordination, impact analysis, and risk-based decision making
- Participate in post-incident reviews, root cause analysis, corrective action planning, and systemic reliability improvements
- Develop and maintain operational runbooks, readiness criteria, triage models, and resilience practices
- Contribute to design and architecture reviews across teams, services, dependencies, and operational domains
- Partner with SRE, platform, product, infrastructure, security, and business stakeholders
- Translate technical concepts, risks, and tradeoffs for technical and non-technical stakeholders
- Mentor engineers through technical leadership, code and design reviews, documentation, and operational coaching
- Participate in a 24x7 on-call rotation supporting incident response and production support for mission-critical platforms and processes
Requirements
What you’ll need- Hands-on proficiency in multiple languages, including Go, Java, Python, or C#
- Experience building production-grade full stack applications on Kubernetes and serverless technologies such as KNative in Azure and AWS
- Experience with SQL and NoSQL technologies and cloud-native services
- Experience building and using data pipelines, analytics, and dashboards using Spark, Trino, Grafana, Superset, or Power BI
- Experience with OpenTelemetry and observability platforms such as Grafana, Datadog, Splunk, or Azure Monitor
- Experience with incident management platforms such as PagerDuty
- Proficiency with AI-assisted development tools such as Claude Code, Cursor, and GitHub Copilot
- Experience improving incident, post-incident review, or reliability processes through automation, data, and cross-team collaboration
- Strong incident forensics and root cause analysis skills
- Strong understanding of observability, reliability engineering, incident management, and post-incident improvement practices
- Experience supporting incident response and high-severity production incidents in complex environments
- Strong software engineering fundamentals and system design skills
- Ability to contribute to technical design and architecture decisions in complex distributed systems
- Strong communication skills and ability to present findings clearly to leadership
- 4+ years of professional software engineering experience
- 3+ years of experience with architecture, design, system reliability, scalability, and technical delivery for production systems
- 2+ years of experience with open-source frameworks, modern engineering practices, or platform technologies
- 2+ years of experience with Azure, AWS, GCP, or another cloud service provider, or equivalent experience in complex hybrid environments
- Demonstrated ownership of production systems operating in 24x7 environments
- Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience
- Ability to influence engineering outcomes within and across teams
- Clear, concise, professional oral and written communication
- Ability to perform effectively under pressure and in stressful situations
- Ability to participate in a 24x7 on-call rotation
- At this time, GEICO will not sponsor a new applicant for employment authorization for this position
Benefits
Comp & perks- Personalized development programs
- Mentorship
- Certification assistance
- Inclusive and collaborative culture
- Competitive pay
- Benefits
- Flexibility to support your well-being and future
- Equal employment opportunity
- Reasonable accommodations for qualified individuals with disabilities
