FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Principal Site Reliability Engineer
Akamai Technologies. Shape compute-platform strategy and qualification criteria for x64, ARM, accelerator, and inference hardware across firmware, OS, runtime, and reliability .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Linux systems and server platforms, focusing on bare-metal provisioning, infrastructure automation, and incident root-cause investigation. Proficient in enhancing reliability and observability through effective telemetry, diagnostics, and production readiness strategies.
Highest-signal resume keywords
Linux Systems ExpertiseBare-Metal ProvisioningPython ProgrammingObservability and MetricsInfrastructure Automation
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Linux SystemsBare-Metal ProvisioningConfiguration ManagementPythonBashIncident Root-Cause InvestigationKVMQEMULibvirtFirmware Diagnostics
Soft Skills
MentoringDecision-MakingCollaboration
Tools & Technologies
CI/CDTelemetryDiagnosticsAPIsDeployment Pipelines
Industry Keywords
X64ARMAcceleratorsInference HardwareServer LifecyclesBIOS/UEFIBMCsVirtualization TechnologiesVM NetworkingStorage
Tech Stack
Tools & technologiesLinuxPython
About the role
Key responsibilities & impact- Shape compute-platform strategy and qualification criteria for x64, ARM, accelerator, and inference hardware across firmware, OS, runtime, and reliability
- Lead provisioning and CI/CD improvements, including reproducible, safe, and scalable bare-metal provisioning, imaging, configuration, and infrastructure delivery pipelines
- Drive reliability and observability by addressing systemic failures, enhancing telemetry and diagnostics, and establishing safe production rollout and recovery safeguards
- Investigate complex hardware/software incidents, identify root causes, guide decision-making, and convert findings into engineering fixes
- Review designs, mentor engineers, establish standards, and resolve multi-domain issues across teams and vendors
- Enable new hardware through provisioning services and provide seamless firmware upgrade models
- Ensure reliable server behavior across Akamai datacenters
Requirements
What you’ll need- Expertise in Linux systems and server platforms, including boot processes, storage, networking, hardware diagnostics, firmware, BIOS/UEFI, BMCs, and server lifecycles
- Experience designing and operating bare-metal provisioning, configuration management, and scalable infrastructure automation
- Fluency in Python and Bash
- Experience building automation, APIs, deployment pipelines, and debugging multi-layer failures
- Ability to identify whether hardware-test failures originate from firmware, BIOS/UEFI configuration, device firmware, kernel/driver behavior, or the physical test environment
- Expertise in observability, metrics, capacity analysis, incident root-cause investigation, and production readiness
- Practical knowledge of x86 and ARM platforms, accelerators, and inference infrastructure, including drivers, runtimes, and compatibility
- Hands-on experience with Linux virtualization technologies, including KVM, QEMU, and libvirt
- Understanding of nested virtualization, CPU virtualization extensions, VM networking and storage, and Linux-level troubleshooting of virtualized environments
Benefits
Comp & perks- Health, well-being, finances, and life beyond work benefits
- FlexBase workplace flexibility: work at home, in an office, or a combination of both