FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Site Reliability Engineer – Observability Platform
Ford Motor Company. Design and implement scalable observability pipelines spanning metrics, logging, tracing, and alerting .
Posted 10/7/2026full-timeRemote • Washington • United StatesMid-LevelSenior💰 $85,400 - $192,900 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and implementing scalable observability pipelines, utilizing infrastructure-as-code for observability instrumentation, and applying automation to enhance application resilience and reliability. Proficient in cloud-native systems and observability best practices to optimize performance and mitigate stability risks.
Highest-signal resume keywords
SRE ExperienceInfrastructure-As-Code (Terraform, ToFu)Programming (Python, Go, Java/Scala, C, C++)APM and Monitoring Tools (Dynatrace, New Relic, ELK, Splunk, Prometheus)Cloud Platforms (GCP, AWS, Azure)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Observability PipelinesAutomationDestructive TestingPerformance AnalysisError BudgetsCI/CD PipelinesRESTful APIsMicroservicesSoftware DevelopmentAgile Methodologies
Soft Skills
Technical GuidanceCollaborationProblem Solving
Tools & Technologies
DockerKubernetesKafkaDataDogSensuNagios
Industry Keywords
Cloud-Native SystemsObservability StrategiesMTTD/MTTRNoSQL/SQL DatastoresJ2EE
Tech Stack
Tools & technologiesAWSAzureCloudDockerGoogle Cloud PlatformJ2EEJavaKafkaKubernetesMicroservicesNoSQLPrometheusPythonScalaSplunkSpringSpring BootSpringBootSQLTCP/IPTerraformC++Go
About the role
Key responsibilities & impact- Design and implement scalable observability pipelines spanning metrics, logging, tracing, and alerting
- Define and operationalize SLIs and SLOs and establish error budgets
- Build reusable infrastructure-as-code templates and frameworks for observability instrumentation and onboarding
- Architect, design, and develop automation to improve application resilience, recoverability, availability, and scalability
- Perform destructive testing to discover vulnerabilities
- Develop tooling to improve reliability, quality, and time-to-market
- Reduce or eliminate toil through automation
- Collaborate with development teams to build and operate scalable, resilient, cloud-native systems
- Identify stability risks and establish mitigation plans with engineering leadership
- Review technical metrics including errors, response times, caching, capacity, and resource utilization
- Conduct performance analysis and optimization of new and production systems
- Solve complex architecture, design, and business problems by simplifying processes and removing bottlenecks
- Evaluate and integrate emerging technologies and architectures
- Troubleshoot distributed production systems and drive root-cause analysis for platform incidents
- Participate in incident response, support, recovery, and postmortem analysis
- Provide technical guidance and mentorship
- Integrate AI/ML capabilities to enhance anomaly detection, alerting precision, and performance insights
- Embed observability best practices into system design and deployment workflows
Requirements
What you’ll need- Bachelor’s Degree in Computer Science or equivalent experience
- 3+ years of experience in an SRE role
- 5+ years of programming experience with one or more of: Python, Go, Java/Scala, C, or C++
- 3+ years of experience building reusable infrastructure-as-code templates and frameworks in Terraform or ToFu
- 3+ years of experience with APM and monitoring tools such as Dynatrace, New Relic, ELK, Splunk, Prometheus, Sensu, Nagios, Kafka, or DataDog
- 3+ years of experience with J2EE, NoSQL/SQL datastores, Spring Boot, GCP/AWS/Azure, and Docker/Kubernetes in developing multi-tier applications
- Experience with RESTful APIs and microservices platforms
- Working knowledge of the TCP/IP stack, internet routing, and load balancing
- Strong proficiency with Google Cloud Platform and its library of services
- Experience with automated, test-driven development in CI/CD pipelines
- Thorough understanding of software development and agile methodologies
- Understanding of, and ability to implement, effective observability strategies to improve MTTD/MTTR
- Must be legally authorized to work in the United States
- Visa sponsorship is not available for this position
Benefits
Comp & perks- Immediate medical, dental, vision and prescription drug coverage
- Flexible family care days
- Paid parental leave
- New parent ramp-up programs
- Subsidized back-up childcare
- Family building benefits including adoption and surrogacy expense reimbursement and fertility treatments
- Vehicle discount program for employees and family members and management leases
- Tuition assistance
- Established and active employee resource groups
- Paid time off for individual and team community service
- Generous schedule of paid holidays, including the week between Christmas and New Year’s Day
- Paid time off and the option to purchase additional vacation time