FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in architecting, scaling, and optimizing SLURM environments for high-performance computing and AI workloads, with a strong focus on workload scheduling, resource optimization, and multi-tenant cluster management. Proven ability to lead technical initiatives and mentor engineering teams while integrating modern containerization and orchestration technologies.
Highest-signal resume keywords
SLURM AdministrationWorkload SchedulingKubernetes FundamentalsInfrastructure-as-Code (Ansible, Terraform)GPU Scheduling (NVIDIA MIG, MPI, NCCL)
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Workload SchedulingSLURM AdministrationMulti-Tenant Cluster OptimizationScheduler TuningDatabase Performance ManagementTelemetry AnalysisContainerization (Singularity, Apptainer)Infrastructure-as-Code (Ansible, Terraform)GPU ManagementHigh-Speed Interconnects (InfiniBand, RoCE)
Soft Skills
LeadershipMentoringCollaborationProblem-SolvingInitiative Coordination
Tools & Technologies
SLURMKubernetesDockerSingularityApptainerEnrootTerraformAnsibleMPINCCL
Industry Keywords
High-Performance ComputingAI WorkflowsPharmaceutical R&DScientific ResearchObservability
Tech Stack
Tools & technologiesAnsibleCloudDockerKubernetesNode.jsTerraform
About the role
Key responsibilities & impact- Own and advance the scheduler technology stack across Roche’s High-Performance Computing platforms
- Drive efficient scheduling, policy management, and resource optimization of multi-node CPU and GPU environments
- Bridge traditional scientific computing with modern AI paradigms
- Solve scheduling and infrastructure challenges impacting Roche’s compute architecture
- Enable researchers, data scientists, and engineers to execute compute workloads reliably and efficiently
- Architect, scale, and maintain SLURM across heterogeneous HPC and AI environments
- Design and tune advanced SLURM configurations, including custom plugins, topology-aware scheduling, GRES/GPU management, dynamic priority trees, and QoS/fair-share policies
- Evaluate and implement integrations between SLURM, Kubernetes, and orchestration platforms such as SLURM Slinky or Run:ai
- Integrate Singularity/Apptainer containerization across SLURM and support hybrid AI/HPC workloads
- Solve multi-tenant bottlenecks including GPU allocation overhead, MPI/NCCL communication failures, and complex workload failures
- Lead global cross-functional initiatives to establish workload orchestration standards, policies, and architectural patterns
- Mentor junior and mid-level engineers and drive continuous learning
- Partner with Observability Engineers to establish telemetry dashboards for SLURM job efficiency, queue wait times, and hardware utilization
- Use configuration-as-code to deploy policies uniformly
Requirements
What you’ll need- Bachelor’s or advanced degree in Computer Science, Applied Mathematics, Computational Engineering, or a related technical discipline
- Extensive systems engineering experience with deep specialization in workload scheduling, SLURM administration, and multi-tenant cluster optimization
- Demonstrated track record of leading complex technical initiatives and mentoring engineering peers
- Proven experience in life sciences, pharmaceutical R&D, or high-performance scientific research environments
- Subject matter expertise in architecting, scaling, upgrading, and optimizing production SLURM environments, including scheduler/backfill tuning, partition and topology design, priority/fair-share/QoS policies, cgroups, HA architecture, plugin integration, and GRES/TRES modeling for GPUs and specialized resources
- Deep expertise in SlurmDBD and accounting architecture, database performance and lifecycle management, scheduler telemetry and health monitoring, workload efficiency analysis, queue/wait-time diagnostics, utilization analysis, and troubleshooting complex controller, database, node, and workload interactions
- Hands-on experience with Kubernetes fundamentals and container runtimes (Singularity, Apptainer, Enroot, Docker) within an HPC context
- Deep familiarity with GPU scheduling (NVIDIA MIG, fractionalization), high-speed interconnects (InfiniBand, RoCE), and multi-node communication frameworks (MPI, NCCL)
- Advanced proficiency with Infrastructure-as-Code (Ansible, Terraform) to automate scheduler deployments, configuration drift management, and telemetry pipelines
- Broad knowledge across HPC, AI infrastructure, Kubernetes, containers, networking/interconnects, observability, automation, and capacity management
- Proven ability to troubleshoot complex, unprecedented failure modes at the intersection of hardware, OS, schedulers, and workloads
- Strong leadership presence with dedication to mentoring colleagues, driving technical standards, and collaborating with global cross-functional teams
- Demonstrated ability to coordinate initiatives across diverse global business units, IT functions, and scientific research stakeholders
- Passion for guiding the convergence of traditional HPC schedulers like SLURM with cloud-native, Kubernetes-driven AI workflows
