FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Staff Engineer, Site Reliability Engineering
General Motors. Design and implement scalable, fault-tolerant, observable infrastructure for vehicle telemetry, data ingestion, and platform operations .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and implementing scalable, fault-tolerant infrastructure for vehicle telemetry and data ingestion, with a strong focus on CI/CD pipelines, observability, and incident management. Proven ability to lead teams, mentor engineers, and influence technical direction while balancing reliability, performance, and cost.
Highest-signal resume keywords
SRE ExperienceCI/CD Pipeline DesignCloud-Native Systems (Azure, AWS, GCP)Programming in Python, Go, or JavaObservability Patterns and Incident Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
CI/CD Pipeline DesignCloud-Native SystemsProgramming in PythonProgramming in GoProgramming in JavaObservability PatternsIncident ManagementService Level Objectives (SLOs)Service Level Indicators (SLIs)Automated Incident Response
Soft Skills
MentoringCommunication SkillsInfluencing Without AuthorityWorking Under PressureAccountability
Tools & Technologies
Azure DatabricksPrometheusGrafanaOpenTelemetryTerraformGitHub ActionsArgo CDKafkaPulsarFivetran
Certifications & Qualifications
BS in Computer ScienceMS in EngineeringPhD in PhysicsPhD in Mathematics
Industry Keywords
Vehicle TelemetryConnected-Vehicle PlatformsHigh-Volume Event-Driven SystemsAI-Assisted Software DevelopmentContinuous Reliability Improvement
Tech Stack
Tools & technologiesApacheAWSAzureCloudGoogle Cloud PlatformGrafanaJavaKafkaPrometheusPulsarPythonTerraformGo
About the role
Key responsibilities & impact- Design and implement scalable, fault-tolerant, observable infrastructure for vehicle telemetry, data ingestion, and platform operations
- Lead production readiness efforts across multiple teams
- Shape reliability standards and guide architectural improvements
- Design, implement, and improve CI/CD delivery pipelines
- Establish quality gates, artifact promotion, deployment verification, progressive delivery, and rollback practices
- Define and implement SLOs, SLIs, observability, monitoring, alerting, runbooks, and operational best practices
- Build and improve reusable AI workflows, skills, and evaluations
- Automate incident intake, triage, diagnostics, remediation, evidence collection, service requests, and customer-facing status workflows
- Participate in weekly on-call rotations with 12-hour shifts cycling every eight weeks
- Lead incident response and coordinate mitigation and recovery
- Conduct post-incident reviews and drive durable system-level fixes
- Partner with internal customers to understand needs, explain trade-offs, and improve service outcomes
- Influence technical direction, mentor engineers, and raise engineering standards through reviews, documentation, and hands-on leadership
- Balance reliability, performance, security, delivery speed, and cost in technical decisions
Requirements
What you’ll need- 8+ years in SRE, DevOps, or systems engineering
- Experience managing or mentoring high-impact teams
- Experience designing, building, and maintaining high-scale, cloud-native production systems, preferably Azure, AWS, or GCP
- Hands-on experience architecting observability patterns, including standardized instrumentation, OTEL collector configuration, SLO/SLI definitions, monitors, alerts, and dashboards
- Strong understanding of production readiness, service ownership, SLOs, incident management, post-incident learning, and continuous reliability improvement
- Experience participating in on-call rotations and leading technical responses to production incidents
- Experience designing, operating, and improving CI/CD pipelines; understanding of GitOps, release strategies, quality gates, deployment verification, progressive delivery, and safe rollback
- Strong programming ability in Python, Go, Java, or a comparable language
- Disciplined code review, version control, testing, and maintainability practices
- Familiarity with AI-assisted software development, LLM application practices, agentic workflows, and evaluation techniques
- Ability to work effectively with internal customers in difficult or high-pressure situations
- Ability to operate with pace, judgment, and accountability
- Ability to influence without formal authority and increase adoption of shared engineering patterns
- Strong written and verbal communication skills
- BS / MS / PhD in computer science, engineering, physics, mathematics, or another relevant technical field
- Preferred experience with Azure Databricks, Azure Event Hubs, AKS, Helm, Kustomize, Terraform, GitHub Actions, Argo CD, GitOps, Prometheus, Grafana, Datadog, OpenTelemetry, Promptfoo, CoPilot, Fivetran, Apache Flink, Kafka, Pulsar, vehicle telemetry, connected-vehicle platforms, or high-volume event-driven systems
- Must not require GM immigration sponsorship now or in the future, including H1-B, OPT, STEM OPT, CPT, TN, or J-1 sponsorship
Benefits
Comp & perks- Benefits supporting well-being at work and home
- Total Rewards resources
- Reasonable accommodations for applicants with disabilities
- No relocation benefits; relocation costs are the selected candidate’s responsibility