Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Mozn

Distributed Systems Engineer III

Mozn

. Build, operate, and continuously improve production cloud-native platforms running distributed workloads .

Posted 9/21/2026full-timeRemote • EgyptMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and operating cloud-native platforms with a strong focus on Kubernetes and Apache Kafka. Proficient in automation, infrastructure-as-code, and ensuring high availability and resilience in distributed systems.

Highest-signal resume keywords
Kubernetes OperationsApache Kafka ManagementMySQL/PostgreSQL ExpertiseInfrastructure-as-Code with TerraformProduction Troubleshooting Skills

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesApache KafkaMySQLPostgreSQLTerraformPythonBashGoArgoCDDistributed Systems
Soft Skills
Problem-SolvingCollaborationIncident Resolution
Tools & Technologies
PrometheusGrafanaOpenSearchELKGitOpsCDC PlatformsStarRocksClickHouseApache DorisFlink
Industry Keywords
Cloud-NativeInfrastructure EngineeringPlatform EngineeringSREData Infrastructure

Tech Stack

Tools & technologies
ApacheAWSAzureCloudDistributed SystemsGoogle Cloud PlatformGrafanaJavaKafkaKubernetesMySQLNode.jsPostgresPrometheusPythonSparkTerraformGo

About the role

Key responsibilities & impact
  • Build, operate, and continuously improve production cloud-native platforms running distributed workloads
  • Work hands-on with Kubernetes, including upgrades, node pools, workload lifecycle, troubleshooting, and platform operations
  • Deploy and manage workloads using ArgoCD, Helm, GitOps, Terraform, and automation
  • Design and operate systems focused on scalability, availability, resilience, performance, and operational simplicity
  • Troubleshoot complex issues across Kubernetes, cloud infrastructure, networking, storage, applications, and distributed services
  • Operate and troubleshoot Apache Kafka in production across high-throughput and distributed workloads
  • Manage Kafka topics, partitions, replication, consumer groups, retention, throughput, latency, and failure recovery
  • Integrate Kafka with databases and applications using Kafka Connect, Debezium, or similar CDC/event-streaming platforms
  • Operate and troubleshoot MySQL and/or PostgreSQL, including replication, high availability, backup, recovery, performance, and migrations
  • Support data-intensive workloads and analytical platforms such as StarRocks, ClickHouse, Apache Doris, or similar technologies
  • Design and operate resilient platforms across node, service, zone, and infrastructure failures
  • Implement and validate backup, recovery, disaster recovery, failover, and business-continuity capabilities
  • Apply multi-zone, multi-region, and active-active architecture principles where appropriate
  • Build multi-tenant platforms with appropriate isolation, scalability, resource management, and reliability
  • Participate in disaster-recovery exercises, failure simulations, migrations, and resilience initiatives
  • Automate infrastructure and platform lifecycle operations using Terraform, Python, Bash, Go, or similar technologies
  • Build reliable deployment and GitOps workflows and reduce manual operational effort
  • Implement monitoring, logging, alerting, and observability for distributed workloads
  • Participate in production incident response, root-cause analysis, and long-term reliability improvements
  • Contribute infrastructure for AI, machine-learning, and data-intensive workloads
  • Evolve cloud and Kubernetes platforms for AI workloads, data pipelines, model-serving infrastructure, and platform services
  • Address compute, GPU, networking, storage, data movement, observability, and workload-isolation requirements for AI platforms
  • Build reusable platform capabilities with engineering teams to run AI and data workloads reliably at scale
  • Stay current with infrastructure patterns across AI platforms, distributed data systems, and cloud-native technologies

Requirements

What you’ll need
  • 4–7 years of experience in Platform Engineering, Infrastructure Engineering, Distributed Systems, SRE, Backend Engineering, Data Infrastructure, or a related field
  • Strong production experience with Apache Kafka — mandatory
  • Hands-on production experience with Kubernetes — mandatory
  • Strong experience with at least one of MySQL or PostgreSQL — mandatory
  • Solid understanding of distributed-systems fundamentals including replication, partitioning, consistency, availability, fault tolerance, scalability, and failure recovery
  • Experience with ArgoCD/GitOps and infrastructure-as-code such as Terraform
  • Experience operating workloads on a public cloud such as GCP, OCI, AWS, or Azure
  • Strong production troubleshooting and incident-resolution skills
  • Experience with automation or scripting using Python, Bash, Go, Java, or similar languages
  • Understanding of high-availability, disaster-recovery, and multi-tenant architecture fundamentals
  • Experience with observability and operational tooling such as Prometheus, Grafana, OpenSearch/ELK, LGTM, or equivalent
  • Strong understanding of infrastructure and networking fundamentals in cloud-native environments
  • Good-to-have: experience with Kafka Connect, Debezium, Kafka Streams, or CDC platforms
  • Good-to-have: experience with distributed analytical databases such as StarRocks, ClickHouse, Apache Doris, or similar
  • Good-to-have: experience with Flink, Spark, or other distributed data-processing systems
  • Good-to-have: experience executing large-scale data, database, application, or infrastructure migrations
  • Good-to-have: experience with active-active, multi-zone, or multi-region systems
  • Good-to-have: experience operating stateful workloads on Kubernetes
  • Good-to-have: experience supporting AI/ML infrastructure or GPU-based workloads
  • Good-to-have: experience with cloud networking, service mesh, ingress, load balancing, or storage platforms
  • Good-to-have: contributions to Kubernetes, Kafka, distributed-systems, or other open-source infrastructure projects

Benefits

Comp & perks
  • Competitive compensation
  • Top-tier health insurance
  • Enabling culture
  • Responsibility and trust
  • Freedom and autonomy in decision-making
  • Fun and dynamic workplace
  • Opportunity to work alongside leading AI professionals
  • Inclusive and diverse workplace culture