FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating cloud-native platforms with a strong focus on Kubernetes and Apache Kafka. Proficient in automation, infrastructure-as-code, and ensuring high availability and resilience in distributed systems.
Highest-signal resume keywords
Kubernetes OperationsApache Kafka ManagementMySQL/PostgreSQL ExpertiseInfrastructure-as-Code with TerraformProduction Troubleshooting Skills
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesApache KafkaMySQLPostgreSQLTerraformPythonBashGoArgoCDDistributed Systems
Soft Skills
Problem-SolvingCollaborationIncident Resolution
Tools & Technologies
PrometheusGrafanaOpenSearchELKGitOpsCDC PlatformsStarRocksClickHouseApache DorisFlink
Industry Keywords
Cloud-NativeInfrastructure EngineeringPlatform EngineeringSREData Infrastructure
Tech Stack
Tools & technologiesApacheAWSAzureCloudDistributed SystemsGoogle Cloud PlatformGrafanaJavaKafkaKubernetesMySQLNode.jsPostgresPrometheusPythonSparkTerraformGo
About the role
Key responsibilities & impact- Build, operate, and continuously improve production cloud-native platforms running distributed workloads
- Work hands-on with Kubernetes, including upgrades, node pools, workload lifecycle, troubleshooting, and platform operations
- Deploy and manage workloads using ArgoCD, Helm, GitOps, Terraform, and automation
- Design and operate systems focused on scalability, availability, resilience, performance, and operational simplicity
- Troubleshoot complex issues across Kubernetes, cloud infrastructure, networking, storage, applications, and distributed services
- Operate and troubleshoot Apache Kafka in production across high-throughput and distributed workloads
- Manage Kafka topics, partitions, replication, consumer groups, retention, throughput, latency, and failure recovery
- Integrate Kafka with databases and applications using Kafka Connect, Debezium, or similar CDC/event-streaming platforms
- Operate and troubleshoot MySQL and/or PostgreSQL, including replication, high availability, backup, recovery, performance, and migrations
- Support data-intensive workloads and analytical platforms such as StarRocks, ClickHouse, Apache Doris, or similar technologies
- Design and operate resilient platforms across node, service, zone, and infrastructure failures
- Implement and validate backup, recovery, disaster recovery, failover, and business-continuity capabilities
- Apply multi-zone, multi-region, and active-active architecture principles where appropriate
- Build multi-tenant platforms with appropriate isolation, scalability, resource management, and reliability
- Participate in disaster-recovery exercises, failure simulations, migrations, and resilience initiatives
- Automate infrastructure and platform lifecycle operations using Terraform, Python, Bash, Go, or similar technologies
- Build reliable deployment and GitOps workflows and reduce manual operational effort
- Implement monitoring, logging, alerting, and observability for distributed workloads
- Participate in production incident response, root-cause analysis, and long-term reliability improvements
- Contribute infrastructure for AI, machine-learning, and data-intensive workloads
- Evolve cloud and Kubernetes platforms for AI workloads, data pipelines, model-serving infrastructure, and platform services
- Address compute, GPU, networking, storage, data movement, observability, and workload-isolation requirements for AI platforms
- Build reusable platform capabilities with engineering teams to run AI and data workloads reliably at scale
- Stay current with infrastructure patterns across AI platforms, distributed data systems, and cloud-native technologies
Requirements
What you’ll need- 4–7 years of experience in Platform Engineering, Infrastructure Engineering, Distributed Systems, SRE, Backend Engineering, Data Infrastructure, or a related field
- Strong production experience with Apache Kafka — mandatory
- Hands-on production experience with Kubernetes — mandatory
- Strong experience with at least one of MySQL or PostgreSQL — mandatory
- Solid understanding of distributed-systems fundamentals including replication, partitioning, consistency, availability, fault tolerance, scalability, and failure recovery
- Experience with ArgoCD/GitOps and infrastructure-as-code such as Terraform
- Experience operating workloads on a public cloud such as GCP, OCI, AWS, or Azure
- Strong production troubleshooting and incident-resolution skills
- Experience with automation or scripting using Python, Bash, Go, Java, or similar languages
- Understanding of high-availability, disaster-recovery, and multi-tenant architecture fundamentals
- Experience with observability and operational tooling such as Prometheus, Grafana, OpenSearch/ELK, LGTM, or equivalent
- Strong understanding of infrastructure and networking fundamentals in cloud-native environments
- Good-to-have: experience with Kafka Connect, Debezium, Kafka Streams, or CDC platforms
- Good-to-have: experience with distributed analytical databases such as StarRocks, ClickHouse, Apache Doris, or similar
- Good-to-have: experience with Flink, Spark, or other distributed data-processing systems
- Good-to-have: experience executing large-scale data, database, application, or infrastructure migrations
- Good-to-have: experience with active-active, multi-zone, or multi-region systems
- Good-to-have: experience operating stateful workloads on Kubernetes
- Good-to-have: experience supporting AI/ML infrastructure or GPU-based workloads
- Good-to-have: experience with cloud networking, service mesh, ingress, load balancing, or storage platforms
- Good-to-have: contributions to Kubernetes, Kafka, distributed-systems, or other open-source infrastructure projects
Benefits
Comp & perks- Competitive compensation
- Top-tier health insurance
- Enabling culture
- Responsibility and trust
- Freedom and autonomy in decision-making
- Fun and dynamic workplace
- Opportunity to work alongside leading AI professionals
- Inclusive and diverse workplace culture
