Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
CloudLinux

Platform Engineer

CloudLinux

. Run and maintain the observability platform, including team onboarding, cost and capacity monitoring, and alerting.

Posted 10/9/2026full-timeRemote • Spain, Portugal, Montenegro, Bulgaria, SerbiaMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in infrastructure and platform reliability engineering, with a strong focus on Linux systems administration, Kubernetes management, and infrastructure as code using tools like Ansible and Terraform. Proficient in monitoring and alerting systems, technical documentation, and effective communication with engineering teams.

Highest-signal resume keywords
Infrastructure As CodeKubernetes ManagementGitLab AdministrationLinux Systems AdministrationPrometheus And Grafana

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Infrastructure As CodeKubernetesLinux Systems AdministrationGitLab CIAnsibleTerraformPrometheusGrafanaPythonGo
Soft Skills
Strong CommunicationInterpersonal Skills
Tools & Technologies
GitLabAI Engineering AssistantsS3-Compatible Object StorageSelf-Hosted SentryKafkaClickHouseRedis
Industry Keywords
Site Reliability EngineeringProduction MonitoringAlerting DesignCost MonitoringCapacity Monitoring

Tech Stack

Tools & technologies
AnsibleAWSGrafanaKafkaKubernetesLinuxPrometheusPythonRedisTerraformGo

About the role

Key responsibilities & impact
  • Run and maintain the observability platform, including team onboarding, cost and capacity monitoring, and alerting.
  • Run GitLab and the CI runner fleet, including upgrades, capacity, access, backups and restore drills.
  • Keep platform services healthy with production monitoring and runbooks.
  • Research, design and deploy requested services from scratch as code, with monitoring, backups and documentation.
  • Handle developer requests involving access, onboarding, pipeline problems, exporters and dashboards, converting recurring requests into self-service.
  • Respond to incidents, mitigate impact, restore services safely, conduct root-cause analyses and post-mortems, and implement prevention or detection improvements.
  • Ship changes as code through reviewed merge requests.
  • Write runbooks, onboarding guides, maintenance notices and status updates for engineers outside the team.
  • Work with AI agents by delegating collection and drafting, reviewing outputs, and recording learnings for the team.

Requirements

What you’ll need
  • Senior-level experience in infrastructure, platform or site reliability engineering, including at least one production service you were responsible for keeping up.
  • Linux systems administration and debugging on bare metal and virtual machines.
  • Kubernetes in production delivered through GitOps, including cluster upgrades you performed yourself.
  • Infrastructure as code using Ansible and Terraform or OpenTofu, with changes reviewed in merge requests.
  • GitLab administration and GitLab CI in production, self-hosted or SaaS; deep experience with another CI system is acceptable if equivalent depth can be shown.
  • Working knowledge of the Prometheus and Grafana ecosystem, including running it for a team, writing alert rules and dashboards, and reading PromQL.
  • Ability to write technical explanations for engineers outside the team, including runbooks, notices and answers to requests.
  • Strong communication and interpersonal skills.
  • Advanced use of AI engineering assistants such as Claude and Codex.
  • English at upper-intermediate level or higher.
  • Nice to have: alerting design, microVM isolation for CI, S3-compatible object storage operations, AWS cost work, self-hosted Sentry or Kafka/ClickHouse/Redis-backed applications, and Python or Go.

Benefits

Comp & perks
  • A focus on professional development.
  • Interesting and challenging projects.
  • Fully remote work with flexible working hours, which allows you to schedule your day and work from any location worldwide.
  • Paid 24 days of vacation per year, 10 days of national holidays, and unlimited sick leaves.
  • Compensation for private medical insurance.
  • Co-working and gym/sports reimbursement.
  • Budget for education.
  • The opportunity to receive a reward for the most innovative idea that the company can patent.