Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
NVIDIA

HPC Operations Engineer

NVIDIA

. Troubleshoot incoming support requests in a large-scale HPC environment .

Posted 9/29/2026full-timeTel Aviv • IsraelJuniorMid-LevelWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in troubleshooting and operating HPC infrastructure, with a strong focus on automation, configuration management, and collaboration with cross-functional teams to enhance chip development processes.

Highest-signal resume keywords
HPC Infrastructure OperationCentOS/RHEL Linux ProficiencyPython and UNIX ScriptingAnsible Configuration ManagementContainer Technologies Understanding

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
HPC TroubleshootingPython ProgrammingUNIX ScriptingAnsibleDockerCluster Configuration ManagementPerl ProgrammingNetworking (TCP/IP, InfiniBand)Distributed Storage Systems (Lustre, GPFS)Job Scheduler Administration (IBM Spectrum LSF, SLURM)
Soft Skills
Problem-SolvingCommunicationTeamwork
Tools & Technologies
FlexLM License ManagementNFSAutomounterLDAPDNS
Industry Keywords
High-Performance ComputingChip DevelopmentDeployment AutomationOperational MonitoringCompute Infrastructure

Tech Stack

Tools & technologies
AnsibleDNSDockerLinuxNFSPerlPythonTCP/IPUnix

About the role

Key responsibilities & impact
  • Troubleshoot incoming support requests in a large-scale HPC environment
  • Improve deployment automation, configuration management, observability, and operational monitoring through automation
  • Operate HPC infrastructure day to day
  • Ensure compute servers run an accurate operating system and configuration
  • Resolve complex issues from bare metal to application level
  • Collaborate with specialist teams to drive issues to closure
  • Collaborate with domain experts to improve infrastructure use in the chip development process
  • Contribute to overall quality and improve time to market for next-generation chips

Requirements

What you’ll need
  • BS in Computer Science or similar degree or equivalent experience
  • 2+ years of experience
  • Proficiency in coordinating CentOS/RHEL Linux distributions
  • Understanding of container technologies such as Docker
  • Proficiency in Python and UNIX scripting languages such as bash
  • Excellent problem-solving skills, with the ability to analyze sophisticated systems, identify bottlenecks, and implement scalable solutions
  • Excellent communication and teamwork skills, with the ability to work optimally with teams with varied strengths and individuals
  • Proven understanding of cluster configuration management tools such as Ansible
  • Understanding of Linux technologies such as NFS, automounter, LDAP, DNS, and TCP/IP networking in Red Hat Linux distributions
  • Familiarity with job scheduler administration such as IBM Spectrum LSF or SLURM
  • Experience building/operating large-scale compute infrastructure
  • Knowledge of the FlexLM license management system
  • Proficiency in Perl for maintaining legacy automation scripts
  • Familiarity with high-speed networking such as InfiniBand, RDMA, and RoCE
  • Familiarity with fast, distributed storage systems such as Lustre and GPFS