Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Roche

Orchestration Workload Engineer – AI Factory

Roche

. Own and advance the scheduler technology stack across Roche’s High-Performance Computing platforms .

Posted 9/18/2026full-timeSouth San Francisco • California • United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in architecting, scaling, and optimizing SLURM environments for high-performance computing and AI workloads, with a strong focus on workload scheduling, resource optimization, and multi-tenant cluster management. Proven ability to lead technical initiatives and mentor engineering teams while integrating modern containerization and orchestration technologies.

Highest-signal resume keywords
SLURM AdministrationWorkload SchedulingKubernetes FundamentalsInfrastructure-as-Code (Ansible, Terraform)GPU Scheduling (NVIDIA MIG, MPI, NCCL)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Workload SchedulingSLURM AdministrationMulti-Tenant Cluster OptimizationScheduler TuningDatabase Performance ManagementTelemetry AnalysisContainerization (Singularity, Apptainer)Infrastructure-as-Code (Ansible, Terraform)GPU ManagementHigh-Speed Interconnects (InfiniBand, RoCE)
Soft Skills
LeadershipMentoringCollaborationProblem-SolvingInitiative Coordination
Tools & Technologies
SLURMKubernetesDockerSingularityApptainerEnrootTerraformAnsibleMPINCCL
Industry Keywords
High-Performance ComputingAI WorkflowsPharmaceutical R&DScientific ResearchObservability

Tech Stack

Tools & technologies
AnsibleCloudDockerKubernetesNode.jsTerraform

About the role

Key responsibilities & impact
  • Own and advance the scheduler technology stack across Roche’s High-Performance Computing platforms
  • Drive efficient scheduling, policy management, and resource optimization of multi-node CPU and GPU environments
  • Bridge traditional scientific computing with modern AI paradigms
  • Solve scheduling and infrastructure challenges impacting Roche’s compute architecture
  • Enable researchers, data scientists, and engineers to execute compute workloads reliably and efficiently
  • Architect, scale, and maintain SLURM across heterogeneous HPC and AI environments
  • Design and tune advanced SLURM configurations, including custom plugins, topology-aware scheduling, GRES/GPU management, dynamic priority trees, and QoS/fair-share policies
  • Evaluate and implement integrations between SLURM, Kubernetes, and orchestration platforms such as SLURM Slinky or Run:ai
  • Integrate Singularity/Apptainer containerization across SLURM and support hybrid AI/HPC workloads
  • Solve multi-tenant bottlenecks including GPU allocation overhead, MPI/NCCL communication failures, and complex workload failures
  • Lead global cross-functional initiatives to establish workload orchestration standards, policies, and architectural patterns
  • Mentor junior and mid-level engineers and drive continuous learning
  • Partner with Observability Engineers to establish telemetry dashboards for SLURM job efficiency, queue wait times, and hardware utilization
  • Use configuration-as-code to deploy policies uniformly

Requirements

What you’ll need
  • Bachelor’s or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a related technical discipline
  • Extensive systems engineering experience with deep specialization in workload scheduling, SLURM administration, and multi-tenant cluster optimization
  • Demonstrated track record of leading complex technical initiatives and mentoring engineering peers
  • Proven experience in life sciences, pharmaceutical R&D, or high-performance scientific research environments
  • Subject matter expertise in architecting, scaling, upgrading, and optimizing production SLURM environments, including scheduler/backfill tuning, partition and topology design, priority/fair-share/QoS policies, cgroups, HA architecture, plugin integration, and GRES/TRES modeling for GPUs and specialized resources
  • Deep expertise in SlurmDBD and accounting architecture, database performance and lifecycle management, scheduler telemetry and health monitoring, workload efficiency analysis, queue/wait-time diagnostics, utilization analysis, and troubleshooting complex controller, database, node, and workload interactions
  • Hands-on experience with Kubernetes fundamentals and container runtimes (Singularity, Apptainer, Enroot, Docker) within an HPC context
  • Deep familiarity with GPU scheduling (NVIDIA MIG, fractionalization), high-speed interconnects (InfiniBand, RoCE), and multi-node communication frameworks (MPI, NCCL)
  • Advanced proficiency with Infrastructure-as-Code (Ansible, Terraform) to automate scheduler deployments, configuration drift management, and telemetry pipelines
  • Broad knowledge across HPC, AI infrastructure, Kubernetes, containers, networking/interconnects, observability, automation, and capacity management
  • Proven ability to troubleshoot complex, unprecedented failure modes at the intersection of hardware, OS, schedulers, and workloads
  • Strong leadership presence with dedication to mentoring colleagues, driving technical standards, and collaborating with global cross-functional teams
  • Demonstrated ability to coordinate initiatives across diverse global business units, IT functions, and scientific research stakeholders
  • Passion for guiding the convergence of traditional HPC schedulers like SLURM with cloud-native, Kubernetes-driven AI workflows