FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Site Reliability Engineer – MAAS
Pragmatike. Operate and maintain large-scale Linux infrastructure across Debian/Ubuntu-based bare-metal and virtualized environments .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expert-level Linux administration, particularly in Debian/Ubuntu, and strong production experience with Kubernetes and MAAS for bare-metal provisioning. Proficient in automation using Ansible, Bash, and Terraform, with a solid understanding of network engineering and infrastructure security practices.
Highest-signal resume keywords
Linux AdministrationKubernetes OperationsMAAS ProvisioningAnsible AutomationNetwork Engineering
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Linux AdministrationKubernetesMAASAnsibleBashPythonTerraformPrometheusGrafanaNetwork Engineering
Soft Skills
Autonomous WorkProblem-SolvingCollaboration
Tools & Technologies
ProxmoxKVM/libvirtOpenStackVMwareIPMI/RedfishBMCs
Industry Keywords
Infrastructure EngineeringSite Reliability EngineeringDistributed SystemsInfrastructure SecurityOperational Efficiency
Tech Stack
Tools & technologiesAnsibleCloudDistributed SystemsDNSFirewallsGrafanaKubernetesLinuxNode.jsOpenStackPrometheusPythonTerraformVMware
About the role
Key responsibilities & impact- Operate and maintain large-scale Linux infrastructure across Debian/Ubuntu-based bare-metal and virtualized environments
- Own MAAS-based bare-metal provisioning, including region/rack controllers, PXE, commissioning, cloud-init, node lifecycle, and API/CLI automation
- Operate and maintain production Kubernetes clusters, including upgrades, node pools, networking, storage, security hardening, and troubleshooting
- Design and maintain multi-site networking across VLANs, L2/L3 routing, bonded interfaces, VPNs, firewalls, and DNS
- Automate infrastructure provisioning and operations using Ansible, Bash/Python, OpenTofu/Terraform, and Git-based workflows
- Build and maintain automated deployment workflows including PXE, Preseed, and cloud-init
- Operate observability platforms using Prometheus, Grafana, Alertmanager, VictoriaMetrics/VictoriaLogs, or comparable tooling
- Define and improve SLIs, SLOs, alerting, and reliability practices across infrastructure and platform services
- Lead infrastructure incident response, troubleshooting, escalation, and post-incident improvements
- Maintain on-call processes and operational coverage across distributed environments
- Work close to the hardware layer, including IPMI/Redfish, BMCs, RAID, storage, hardware diagnostics, and GPU infrastructure
- Manage virtualization platforms including Proxmox, KVM/libvirt, OpenStack, or VMware, including GPU passthrough where required
- Build and maintain internal infrastructure tooling for host discovery, configuration, IPAM, hardware health, and operational automation
- Own infrastructure lifecycle activities including site onboarding, maintenance, decommissioning, drift detection, and operational runbooks
- Work closely with engineering and cross-functional teams to improve reliability, resource utilization, and operational efficiency
Requirements
What you’ll need- 5+ years of hands-on SRE, Infrastructure, Systems, or Platform Engineering experience
- Expert-level Linux administration, particularly Debian/Ubuntu
- Strong production experience with MAAS and bare-metal provisioning
- Expert-level, hands-on experience operating Kubernetes in production, including cluster lifecycle, networking, storage, upgrades, and troubleshooting
- Strong network engineering skills across VLANs, L2/L3 routing, bonding, VPNs, firewalls, and DNS
- Strong automation skills with Ansible, Bash and/or Python
- Experience with Terraform/OpenTofu and Git-based infrastructure workflows
- Production experience with Prometheus/Grafana or comparable observability platforms
- Experience with incident response, on-call operations, monitoring, alerting, and reliability practices
- Experience with Proxmox, KVM/libvirt, OpenStack, VMware, or comparable virtualization technologies
- Experience with bare-metal hardware, BMCs, IPMI/Redfish, storage, and hardware troubleshooting
- Strong understanding of distributed systems, container orchestration, and infrastructure reliability
- Experience with infrastructure security including RBAC, firewalls, network policies, secrets management, and security hardening
- Ability to create SOPs, runbooks, and operational processes from scratch
- Comfortable working autonomously in a fast-paced, engineering-driven environment
- Fluent English required
Benefits
Comp & perks- 100% remote with flexible working hours
- High-impact role with significant technical ownership and autonomy
- Work directly with bare-metal, Kubernetes, networking, and GPU infrastructure
- International, engineering-driven team
- Strong focus on automation, reliability, and infrastructure at scale
- Opportunity to shape the architecture and operational foundations of a growing cloud platform