FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive expertise in managing cloud infrastructure, particularly in Kubernetes and GCP environments, with a strong focus on observability, incident management, and operational readiness. Proficient in authoring runbooks and leading capacity planning while mentoring junior engineers in best practices.
Highest-signal resume keywords
Kubernetes ExpertiseGCP Cloud OperationsPrometheus Observability StackPostgreSQL ManagementInfrastructure Provisioning with Terraform
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesGCPPostgreSQLRedisPrometheusGrafanaTerraformAnsibleGitOpsCloud Infrastructure Engineering
Soft Skills
Strong CommunicationMentoring
Tools & Technologies
GrafanaVictoriaMetricsOpenTelemetryArgoCDHelmLokiELKPub/Sub
Certifications & Qualifications
CKACKSAWS Solutions Architect ProfessionalRed Hat Certified Architect
Industry Keywords
Site Reliability Engineering24x7 EnvironmentCloud OperationsObservabilityCapacity Planning
Tech Stack
Tools & technologiesAnsibleAWSCloudFluxGoogle Cloud PlatformGrafanaKubernetesNode.jsOpenStackPostgresPrometheusPythonRedisTerraformGo
About the role
Key responsibilities & impact- Own 24x7 cloud infrastructure health across Skylo’s hybrid production environment
- Operate GCP public cloud and on-premise private cloud infrastructure
- Monitor and triage infrastructure alarms using OSS dashboards, Grafana/VictoriaMetrics, GCP Cloud Monitoring, and Loki
- Execute runbooks for GKE node recovery, pod eviction/rescheduling, PVC repair, database failover, Prometheus WAL recovery, ArgoCD drift remediation, and certificate rotation
- Own the observability pipeline, including Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, and alert routing
- Maintain PostgreSQL replication, backups, restores, failover testing, query performance, and Redis operations
- Ensure log aggregation pipeline health using Loki or ELK
- Partner with Network Implementation on infrastructure changes and operational readiness
- Serve as L3 escalation authority for Cloud Infrastructure incidents
- Lead troubleshooting bridges and diagnose Kubernetes, storage, network, database, and GitOps failures
- Participate in the global 24x7 on-call rotation
- Define and maintain SLOs, track error budgets, reduce toil, and lead capacity planning
- Own infrastructure root-cause analyses and post-incident action items
- Author and maintain Cloud Infrastructure runbooks and SOPs
- Validate operational readiness for infrastructure expansions, upgrades, and hardware deployments
- Represent Cloud Infrastructure in architecture reviews and collaborate with NRE, security, platform engineering, and automation teams
- Mentor Senior NREs in Kubernetes, storage, database reliability, observability, and escalation practices
- Oversee ArgoCD, Helm, Terraform, and Ansible changes affecting production
Requirements
What you’ll need- 5+ years of infrastructure engineering, Site Reliability Engineering, or cloud operations in a production 24x7 environment
- Direct on-call ownership for Kubernetes-at-scale environments
- Deep Kubernetes expertise, including multi-cluster operations, node pool management, RBAC, network policies, PVCs, CSI drivers, CRD/operator patterns, and production cluster upgrades
- Hands-on experience operating public cloud and on-premise/private cloud infrastructure
- Production observability stack ownership using Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, and Pub/Sub or equivalent
- PostgreSQL streaming replication, backup/restore, failover procedures, and performance tuning
- Redis cluster operations and persistence management
- Production GitOps experience with ArgoCD or Flux CD
- Helm chart authorship and version management
- Terraform or Ansible for infrastructure provisioning
- SRE fundamentals including SLO/SLI/SLA definition, error budget management, toil measurement, capacity planning, and on-call rotation design
- Container runtime debugging, kernel-level performance analysis, storage subsystem troubleshooting, and network packet flow understanding
- Ability to author runbooks executable independently by less-experienced engineers under incident pressure
- Strong written and verbal communication for RCA documents, structured engineering escalations, and MNO-facing infrastructure summaries
- Preferred: telecom or NTN workload experience
- Preferred: Ceph, Rook, or equivalent distributed storage expertise
- Preferred: KubeVirt, Harvester, or OpenStack experience
- Preferred: BGP, VXLAN, EVPN, software-defined networking, and hardware load balancers
- Preferred: Go or Python development
- Preferred: FinOps experience
- Preferred certifications: CKA, CKS, AWS Solutions Architect Professional, or Red Hat Certified Architect
- Legally authorized to work in the U.S.
Benefits
Comp & perks- Stock option-based equity program
- Medical, dental, and vision benefits
- Retirement plan
- Monthly wellness allowance
- Monthly education reimbursement
- Generous time-off policy
- Paid holidays
- Opportunity to temporarily work abroad
- Access to a world-class team across software, hardware, chipsets, telecom, satellite, and network virtualization
- Flexible approach to work
- Inclusive and diverse workplace culture
