Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Kaltura

DevOps – AI Factory

Kaltura

. Join a global DevOps & Platform team responsible for designing and operating distributed AI production systems at very high scale .

Posted 10/4/2026full-timeIsraelJuniorMid-LevelWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing, building, and operating distributed AI production systems at scale, with a strong focus on cloud platforms like AWS and GCP. Proficient in leveraging modern AI coding tools and implementing CI/CD practices to enhance system reliability and performance.

Highest-signal resume keywords
DevOps ExperienceCloud Infrastructure ExpertiseKubernetes and Helm ProficiencyAI Workload ManagementCI/CD and Automation Skills

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
DevOpsDistributed SystemsAI Factory SystemsGPU InferencingInfrastructure as CodeKubernetesHelmCI/CDNetworkingOperating Systems
Soft Skills
Project LeadershipCollaboration
Tools & Technologies
AWSGCPDockerTerraformGitHub ActionsPrometheusGrafanaKubeFlowRayAI Coding Tools
Industry Keywords
AI Production SystemsObservabilityMonitoringCloud Inferencing ServicesAI Evals

Tech Stack

Tools & technologies
AWSCloudDistributed SystemsDNSDockerGoogle Cloud PlatformGrafanaKubernetesLinuxPrometheusRayTCP/IPTerraform

About the role

Key responsibilities & impact
  • Join a global DevOps & Platform team responsible for designing and operating distributed AI production systems at very high scale
  • Own projects end-to-end, from design through implementation to production rollout
  • Design, build, and operate highly available distributed systems at scale on AWS and GCP
  • Work across Kubernetes, Docker, Helm, CI/CD, automation, data, caching, and messaging layers
  • Modernize and consolidate different stacks and architectures into Kaltura's single coherent platform
  • Build and operate infrastructure for AI workloads, including GPU inference and serving platforms
  • Drive system reliability, monitoring, and observability
  • Troubleshoot complex issues across application, infrastructure, networking, operating system, and data layers
  • Collaborate with Platform, Development, Architecture, AI/ML Research, and Product teams
  • Leverage modern AI coding tools to work at high velocity and quality

Requirements

What you’ll need
  • 2+ years of hands-on experience in DevOps / Platform / SRE roles, operating distributed production systems at large scale, with a proven ability to lead projects independently and drive new implementations end-to-end
  • 2+ years of experience working with cloud Inferencing services and NeoClouds
  • 2+ years of hands-on experience building and maintaining AI Factory systems on KubeFlow/Ray/similar
  • 2+ years of hands on experience with GPU Inferencing (vLLM/SGLang/Nvidia Triton)
  • Deep hands-on expertise with multicloud AWS/GCP/NeoClouds (RunPod, TogetherAI, vast.ai) including Infrastructure as Code (Terraform or similar)
  • Deep hands-on expertise with Kubernetes and Helm, including a strong understanding of how Kubernetes internals work and how to operate them in production, alongside Docker / containerized systems
  • Strong knowledge of networking and operating systems (TCP/IP, DNS, load balancing, Linux internals)
  • Strong experience with CI/CD and automation (GitHub Actions), monitoring and observability (metrics, tracing, logging, tools like Prometheus / Grafana)
  • High proficiency with current AI coding tools, including a deep understanding of how they work and how to apply them to drive real productivity
  • Familiarity with AI evals and AI observability systems, such as Opik, OpenEval, LangSmith etc. (nice to have)

Benefits

Comp & perks
  • Hybrid, flexible work environment
  • Extended private health (including mental) insurance
  • Personal and professional development programs
  • Occasional Cross company long weekends