Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Ford Motor Company

Site Reliability Engineer – Observability Platform

Ford Motor Company

. Design and implement scalable observability pipelines spanning metrics, logging, tracing, and alerting .

Posted 10/7/2026full-timeRemote • Washington • United StatesMid-LevelSenior💰 $85,400 - $192,900 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and implementing scalable observability pipelines, utilizing infrastructure-as-code for observability instrumentation, and applying automation to enhance application resilience and reliability. Proficient in cloud-native systems and observability best practices to optimize performance and mitigate stability risks.

Highest-signal resume keywords
SRE ExperienceInfrastructure-As-Code (Terraform, ToFu)Programming (Python, Go, Java/Scala, C, C++)APM and Monitoring Tools (Dynatrace, New Relic, ELK, Splunk, Prometheus)Cloud Platforms (GCP, AWS, Azure)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Observability PipelinesAutomationDestructive TestingPerformance AnalysisError BudgetsCI/CD PipelinesRESTful APIsMicroservicesSoftware DevelopmentAgile Methodologies
Soft Skills
Technical GuidanceCollaborationProblem Solving
Tools & Technologies
DockerKubernetesKafkaDataDogSensuNagios
Industry Keywords
Cloud-Native SystemsObservability StrategiesMTTD/MTTRNoSQL/SQL DatastoresJ2EE

Tech Stack

Tools & technologies
AWSAzureCloudDockerGoogle Cloud PlatformJ2EEJavaKafkaKubernetesMicroservicesNoSQLPrometheusPythonScalaSplunkSpringSpring BootSpringBootSQLTCP/IPTerraformC++Go

About the role

Key responsibilities & impact
  • Design and implement scalable observability pipelines spanning metrics, logging, tracing, and alerting
  • Define and operationalize SLIs and SLOs and establish error budgets
  • Build reusable infrastructure-as-code templates and frameworks for observability instrumentation and onboarding
  • Architect, design, and develop automation to improve application resilience, recoverability, availability, and scalability
  • Perform destructive testing to discover vulnerabilities
  • Develop tooling to improve reliability, quality, and time-to-market
  • Reduce or eliminate toil through automation
  • Collaborate with development teams to build and operate scalable, resilient, cloud-native systems
  • Identify stability risks and establish mitigation plans with engineering leadership
  • Review technical metrics including errors, response times, caching, capacity, and resource utilization
  • Conduct performance analysis and optimization of new and production systems
  • Solve complex architecture, design, and business problems by simplifying processes and removing bottlenecks
  • Evaluate and integrate emerging technologies and architectures
  • Troubleshoot distributed production systems and drive root-cause analysis for platform incidents
  • Participate in incident response, support, recovery, and postmortem analysis
  • Provide technical guidance and mentorship
  • Integrate AI/ML capabilities to enhance anomaly detection, alerting precision, and performance insights
  • Embed observability best practices into system design and deployment workflows

Requirements

What you’ll need
  • Bachelor’s Degree in Computer Science or equivalent experience
  • 3+ years of experience in an SRE role
  • 5+ years of programming experience with one or more of: Python, Go, Java/Scala, C, or C++
  • 3+ years of experience building reusable infrastructure-as-code templates and frameworks in Terraform or ToFu
  • 3+ years of experience with APM and monitoring tools such as Dynatrace, New Relic, ELK, Splunk, Prometheus, Sensu, Nagios, Kafka, or DataDog
  • 3+ years of experience with J2EE, NoSQL/SQL datastores, Spring Boot, GCP/AWS/Azure, and Docker/Kubernetes in developing multi-tier applications
  • Experience with RESTful APIs and microservices platforms
  • Working knowledge of the TCP/IP stack, internet routing, and load balancing
  • Strong proficiency with Google Cloud Platform and its library of services
  • Experience with automated, test-driven development in CI/CD pipelines
  • Thorough understanding of software development and agile methodologies
  • Understanding of, and ability to implement, effective observability strategies to improve MTTD/MTTR
  • Must be legally authorized to work in the United States
  • Visa sponsorship is not available for this position

Benefits

Comp & perks
  • Immediate medical, dental, vision and prescription drug coverage
  • Flexible family care days
  • Paid parental leave
  • New parent ramp-up programs
  • Subsidized back-up childcare
  • Family building benefits including adoption and surrogacy expense reimbursement and fertility treatments
  • Vehicle discount program for employees and family members and management leases
  • Tuition assistance
  • Established and active employee resource groups
  • Paid time off for individual and team community service
  • Generous schedule of paid holidays, including the week between Christmas and New Year’s Day
  • Paid time off and the option to purchase additional vacation time