Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
d-Matrix

Senior Staff Site Reliability Engineer, AI Infrastructure

d-Matrix

. Own reliability and availability across colocation server fleets, on-premises lab clusters, cloud environments, and customer-facing platform services .

Posted 9/30/2026full-timeSanta Clara • California • United StatesSenior💰 $175,000 - $265,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on infrastructure management, automation, and incident response. Proficient in provisioning and configuring cloud and on-premises environments, utilizing Infrastructure as Code (IaC) tools, and ensuring high availability and performance of customer-facing services.

Highest-signal resume keywords
Site Reliability Engineering (SRE)Infrastructure as Code (IaC) with Terraform/AnsibleKubernetes Operational ExperienceLinux Systems KnowledgeIncident Response and Root-Cause Analysis

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Infrastructure ManagementCloud OperationsPython ScriptingBash ScriptingPerformance DiagnosticsMonitoring and AlertingCapacity PlanningAutomationRoot-Cause AnalysisNetworking
Soft Skills
CollaborationProblem-SolvingDocumentation
Tools & Technologies
PrometheusGrafanaDataDogSplunkKubernetesAWSAzureGCPAIOps PlatformsHPC Job Schedulers
Certifications & Qualifications
Bachelor's or Master's in Computer ScienceElectrical Engineering
Industry Keywords
ColocationOn-Premises InfrastructureCloud EnvironmentsInfrastructure AutomationService-Level IndicatorsQuality of Service (QoS)Hybrid CloudLarge-Scale InfrastructureCustomer-Facing ServicesPhysical Hardware

Tech Stack

Tools & technologies
AnsibleAWSAzureCloudGoogle Cloud PlatformGrafanaKubernetesLinuxPrometheusPythonSplunkTerraform

About the role

Key responsibilities & impact
  • Own reliability and availability across colocation server fleets, on-premises lab clusters, cloud environments, and customer-facing platform services
  • Provision and configure servers, operating systems, networking, storage, and hardware from bare metal through auto-scaling Kubernetes environments
  • Lead capacity planning and hardware lifecycle management
  • Track cloud spend to support FinOps and workload-placement decisions
  • Drive provisioning, deployment, and operational changes through Terraform and/or Ansible
  • Contribute to shared infrastructure-as-code modules for global SRE and data center services teams
  • Build automation for host lifecycle management, fleet health checks, auto-remediation, self-service tooling, and networking automation
  • Design and maintain monitoring, alerting, and service-level indicators using Prometheus/Grafana, DataDog, Splunk, or equivalent tools
  • Contribute to AIOps-driven detection workflows
  • Participate in an on-call rotation and resolve incidents from bare metal to application layer
  • Produce root-cause analyses for P0/P1 incidents
  • Support internal and external customer platform services, ensuring QoS and uptime commitments
  • Document operational runbooks
  • Partner with hardware and software teams on CI/CD, QA, HPC workloads, silicon development, and customer deployments

Requirements

What you’ll need
  • Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience)
  • 7+ years in SRE, infrastructure engineering, or systems administration
  • Strong Linux systems knowledge
  • Hands-on colocation or on-premises server infrastructure experience, including networking, storage, systemd, kernel parameters, performance diagnostics, physical hardware, rack networking, and bare-metal provisioning
  • Production IaC experience with Terraform and/or Ansible
  • Kubernetes operational experience, including cluster troubleshooting, workload management, storage, and networking
  • Experience with observability tooling such as Prometheus/Grafana, DataDog, Splunk, or equivalent
  • Experience building dashboards and writing alert rules
  • Production-quality Python and/or Bash scripting
  • Incident response experience, including structured triage, RCA production, and follow-through on action items
  • Legal authorization to work in the United States
  • Preferred: experience with customer-facing infrastructure or platform services
  • Preferred: cloud infrastructure operations across AWS, Azure, or GCP, including hybrid cloud/on-premises environments
  • Preferred: production experience with AIOps platforms, intelligent alerting, anomaly detection, or LLM-assisted diagnostics
  • Preferred: HPC job scheduler experience such as Slurm or LSF
  • Preferred: knowledge of InfiniBand, RoCE, or NVLink
  • Preferred: large-scale infrastructure automation experience

Benefits

Comp & perks
  • Equity
  • Bonus/incentive opportunities based on individual and company performance
  • Medical insurance
  • Dental insurance
  • Vision insurance
  • 401(k)
  • Inclusive rewards plan centered around employee and loved ones’ wellbeing