FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Staff Site Reliability Engineer, AI Infrastructure
d-Matrix. Own reliability and availability across colocation server fleets, on-premises lab clusters, cloud environments, and customer-facing platform services .
Posted 9/30/2026full-timeSanta Clara • California • United StatesSenior💰 $175,000 - $265,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on infrastructure management, automation, and incident response. Proficient in provisioning and configuring cloud and on-premises environments, utilizing Infrastructure as Code (IaC) tools, and ensuring high availability and performance of customer-facing services.
Highest-signal resume keywords
Site Reliability Engineering (SRE)Infrastructure as Code (IaC) with Terraform/AnsibleKubernetes Operational ExperienceLinux Systems KnowledgeIncident Response and Root-Cause Analysis
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Infrastructure ManagementCloud OperationsPython ScriptingBash ScriptingPerformance DiagnosticsMonitoring and AlertingCapacity PlanningAutomationRoot-Cause AnalysisNetworking
Soft Skills
CollaborationProblem-SolvingDocumentation
Tools & Technologies
PrometheusGrafanaDataDogSplunkKubernetesAWSAzureGCPAIOps PlatformsHPC Job Schedulers
Certifications & Qualifications
Bachelor's or Master's in Computer ScienceElectrical Engineering
Industry Keywords
ColocationOn-Premises InfrastructureCloud EnvironmentsInfrastructure AutomationService-Level IndicatorsQuality of Service (QoS)Hybrid CloudLarge-Scale InfrastructureCustomer-Facing ServicesPhysical Hardware
Tech Stack
Tools & technologiesAnsibleAWSAzureCloudGoogle Cloud PlatformGrafanaKubernetesLinuxPrometheusPythonSplunkTerraform
About the role
Key responsibilities & impact- Own reliability and availability across colocation server fleets, on-premises lab clusters, cloud environments, and customer-facing platform services
- Provision and configure servers, operating systems, networking, storage, and hardware from bare metal through auto-scaling Kubernetes environments
- Lead capacity planning and hardware lifecycle management
- Track cloud spend to support FinOps and workload-placement decisions
- Drive provisioning, deployment, and operational changes through Terraform and/or Ansible
- Contribute to shared infrastructure-as-code modules for global SRE and data center services teams
- Build automation for host lifecycle management, fleet health checks, auto-remediation, self-service tooling, and networking automation
- Design and maintain monitoring, alerting, and service-level indicators using Prometheus/Grafana, DataDog, Splunk, or equivalent tools
- Contribute to AIOps-driven detection workflows
- Participate in an on-call rotation and resolve incidents from bare metal to application layer
- Produce root-cause analyses for P0/P1 incidents
- Support internal and external customer platform services, ensuring QoS and uptime commitments
- Document operational runbooks
- Partner with hardware and software teams on CI/CD, QA, HPC workloads, silicon development, and customer deployments
Requirements
What you’ll need- Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience)
- 7+ years in SRE, infrastructure engineering, or systems administration
- Strong Linux systems knowledge
- Hands-on colocation or on-premises server infrastructure experience, including networking, storage, systemd, kernel parameters, performance diagnostics, physical hardware, rack networking, and bare-metal provisioning
- Production IaC experience with Terraform and/or Ansible
- Kubernetes operational experience, including cluster troubleshooting, workload management, storage, and networking
- Experience with observability tooling such as Prometheus/Grafana, DataDog, Splunk, or equivalent
- Experience building dashboards and writing alert rules
- Production-quality Python and/or Bash scripting
- Incident response experience, including structured triage, RCA production, and follow-through on action items
- Legal authorization to work in the United States
- Preferred: experience with customer-facing infrastructure or platform services
- Preferred: cloud infrastructure operations across AWS, Azure, or GCP, including hybrid cloud/on-premises environments
- Preferred: production experience with AIOps platforms, intelligent alerting, anomaly detection, or LLM-assisted diagnostics
- Preferred: HPC job scheduler experience such as Slurm or LSF
- Preferred: knowledge of InfiniBand, RoCE, or NVLink
- Preferred: large-scale infrastructure automation experience
Benefits
Comp & perks- Equity
- Bonus/incentive opportunities based on individual and company performance
- Medical insurance
- Dental insurance
- Vision insurance
- 401(k)
- Inclusive rewards plan centered around employee and loved ones’ wellbeing