FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and operating scalable data and AI infrastructure on Kubernetes and public cloud platforms, with a strong focus on improving system performance, reliability, and cost efficiency. Proficient in collaborating across engineering disciplines to translate requirements into durable platform capabilities.
Highest-signal resume keywords
Kubernetes Environment DesignDistributed Data Infrastructure (Kafka, Spark, Flink)Cloud Platforms (AWS, Google Cloud, Microsoft Azure)Programming (Go, Python, Java)Infrastructure-as-Code Practices
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
KubernetesKafkaSparkFlinkAWSGoogle CloudMicrosoft AzureGoPythonJava
Soft Skills
Clear CommunicationOwnership MindsetCollaboration
Tools & Technologies
AI GatewaysModel-Routing PlatformsInfrastructure-as-CodeObservability ToolsSGLangVLLMNVIDIA Triton Inference Server
Industry Keywords
Data InfrastructureAI InfrastructureDistributed SystemsCloud NetworkingCost Efficiency
Tech Stack
Tools & technologiesAWSAzureCloudDNSJavaKafkaKubernetesPythonSparkGo
About the role
Key responsibilities & impact- Design, build, and operate highly available data and AI infrastructure on Kubernetes and public cloud platforms
- Develop scalable platforms for streaming, batch processing, and real-time data workloads using Kafka, Spark, and Flink
- Build and evolve AI infrastructure, including AI gateways, model-routing layers, traffic management, rate limiting, authentication, observability, and usage controls
- Develop self-service capabilities enabling data, AI, and application teams to deploy and operate workloads safely and independently
- Partner with application, data, machine learning, security, and infrastructure teams to translate emerging requirements into durable platform capabilities
- Improve the scalability, reliability, and efficiency of data and AI infrastructure
- Reduce operational effort for deploying and managing data or AI workloads
- Improve visibility into system performance, reliability, capacity, and cost
- Establish reusable platform capabilities adopted by engineering teams
- Help define the technical direction for next-generation data and AI infrastructure
Requirements
What you’ll need- 5+ years in DevOps, SRE, or platform engineering, owning production systems end to end
- Strong experience designing, operating, and troubleshooting production Kubernetes environments
- Experience building or operating distributed data infrastructure with Kafka, Spark, or Flink, or developing AI infrastructure such as AI gateways or model-routing platforms
- Hands-on experience with AWS, Google Cloud, or Microsoft Azure
- Strong knowledge of cloud and container networking, including DNS, load balancing, ingress, service discovery, TLS, routing, and network security
- Proficiency in Go, Python, or Java, with experience writing maintainable production software
- Solid understanding of distributed-systems concepts, including availability, consistency, fault tolerance, backpressure, and horizontal scalability
- Experience operating critical infrastructure using infrastructure-as-code, automated delivery, and modern observability practices
- Strong debugging skills and ability to work methodically across multiple layers of a complex system
- Clear communication skills and track record of collaborating across engineering disciplines
- Ownership mindset and ability to drive important problems to resolution
- Platform engineering experience, particularly building internal developer platforms or paved-road workflows, is a bonus
- Hands-on experience with SGLang, vLLM, or NVIDIA Triton Inference Server is a bonus
- Knowledge of GPU scheduling, batching, model parallelism, memory management, autoscaling, and inference-performance optimization is a bonus
- Experience improving cost efficiency of large-scale data processing or AI inference workloads is a bonus
- Contributions to infrastructure, data-platform, Kubernetes, or AI-serving open-source projects are a bonus
Benefits
Comp & perks- Annual bonuses may be offered
- RSUs or other additional compensation may be offered
- Medical, dental, and vision insurance for US-based employees
- 401(k) plan
- Short-term and long-term disability coverage
- Basic life insurance
- Well-being benefits
- 20 paid days of vacation per calendar year
- 12 paid days of company holidays per calendar year
- Statutory minimum paid annual leave and public holidays
