Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Pragmatike

Senior Site Reliability Engineer – MAAS

Pragmatike

. Operate and maintain large-scale Linux infrastructure across Debian/Ubuntu-based bare-metal and virtualized environments .

Posted 9/29/2026full-timeRemote • ArmeniaSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expert-level Linux administration, particularly in Debian/Ubuntu, and strong production experience with Kubernetes and MAAS for bare-metal provisioning. Proficient in automation using Ansible, Bash, and Terraform, with a solid understanding of network engineering and infrastructure security practices.

Highest-signal resume keywords
Linux AdministrationKubernetes OperationsMAAS ProvisioningAnsible AutomationNetwork Engineering

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Linux AdministrationKubernetesMAASAnsibleBashPythonTerraformPrometheusGrafanaNetwork Engineering
Soft Skills
Autonomous WorkProblem-SolvingCollaboration
Tools & Technologies
ProxmoxKVM/libvirtOpenStackVMwareIPMI/RedfishBMCs
Industry Keywords
Infrastructure EngineeringSite Reliability EngineeringDistributed SystemsInfrastructure SecurityOperational Efficiency

Tech Stack

Tools & technologies
AnsibleCloudDistributed SystemsDNSFirewallsGrafanaKubernetesLinuxNode.jsOpenStackPrometheusPythonTerraformVMware

About the role

Key responsibilities & impact
  • Operate and maintain large-scale Linux infrastructure across Debian/Ubuntu-based bare-metal and virtualized environments
  • Own MAAS-based bare-metal provisioning, including region/rack controllers, PXE, commissioning, cloud-init, node lifecycle, and API/CLI automation
  • Operate and maintain production Kubernetes clusters, including upgrades, node pools, networking, storage, security hardening, and troubleshooting
  • Design and maintain multi-site networking across VLANs, L2/L3 routing, bonded interfaces, VPNs, firewalls, and DNS
  • Automate infrastructure provisioning and operations using Ansible, Bash/Python, OpenTofu/Terraform, and Git-based workflows
  • Build and maintain automated deployment workflows including PXE, Preseed, and cloud-init
  • Operate observability platforms using Prometheus, Grafana, Alertmanager, VictoriaMetrics/VictoriaLogs, or comparable tooling
  • Define and improve SLIs, SLOs, alerting, and reliability practices across infrastructure and platform services
  • Lead infrastructure incident response, troubleshooting, escalation, and post-incident improvements
  • Maintain on-call processes and operational coverage across distributed environments
  • Work close to the hardware layer, including IPMI/Redfish, BMCs, RAID, storage, hardware diagnostics, and GPU infrastructure
  • Manage virtualization platforms including Proxmox, KVM/libvirt, OpenStack, or VMware, including GPU passthrough where required
  • Build and maintain internal infrastructure tooling for host discovery, configuration, IPAM, hardware health, and operational automation
  • Own infrastructure lifecycle activities including site onboarding, maintenance, decommissioning, drift detection, and operational runbooks
  • Work closely with engineering and cross-functional teams to improve reliability, resource utilization, and operational efficiency

Requirements

What you’ll need
  • 5+ years of hands-on SRE, Infrastructure, Systems, or Platform Engineering experience
  • Expert-level Linux administration, particularly Debian/Ubuntu
  • Strong production experience with MAAS and bare-metal provisioning
  • Expert-level, hands-on experience operating Kubernetes in production, including cluster lifecycle, networking, storage, upgrades, and troubleshooting
  • Strong network engineering skills across VLANs, L2/L3 routing, bonding, VPNs, firewalls, and DNS
  • Strong automation skills with Ansible, Bash and/or Python
  • Experience with Terraform/OpenTofu and Git-based infrastructure workflows
  • Production experience with Prometheus/Grafana or comparable observability platforms
  • Experience with incident response, on-call operations, monitoring, alerting, and reliability practices
  • Experience with Proxmox, KVM/libvirt, OpenStack, VMware, or comparable virtualization technologies
  • Experience with bare-metal hardware, BMCs, IPMI/Redfish, storage, and hardware troubleshooting
  • Strong understanding of distributed systems, container orchestration, and infrastructure reliability
  • Experience with infrastructure security including RBAC, firewalls, network policies, secrets management, and security hardening
  • Ability to create SOPs, runbooks, and operational processes from scratch
  • Comfortable working autonomously in a fast-paced, engineering-driven environment
  • Fluent English required

Benefits

Comp & perks
  • 100% remote with flexible working hours
  • High-impact role with significant technical ownership and autonomy
  • Work directly with bare-metal, Kubernetes, networking, and GPU infrastructure
  • International, engineering-driven team
  • Strong focus on automation, reliability, and infrastructure at scale
  • Opportunity to shape the architecture and operational foundations of a growing cloud platform