Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
WEKA

Data Platform Administrator

WEKA

. Run the WEKA environment day to day, including monitoring, health, upgrades, capacity, configuration, and audits .

Posted 10/8/2026full-timeBengaluru • IndiaSeniorLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive experience in administering enterprise infrastructure with a focus on storage solutions, Linux administration, and high-speed networking. Proficient in automation and monitoring tools, with a strong ability to collaborate across teams and communicate effectively with stakeholders.

Highest-signal resume keywords
Enterprise Infrastructure AdministrationLinux Administration (RHEL, Ubuntu)High-Speed Networking (Ethernet, InfiniBand)Kubernetes and Container ManagementPython, Bash, or Ansible Proficiency

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Storage Management (File, Object, Block)Monitoring and Alerting (Prometheus, Grafana)Ticketing Systems (Jira)GPU and HPC Infrastructure ManagementWEKA Environment Management
Soft Skills
Clear Written and Verbal CommunicationStrong Technical Writing
Tools & Technologies
WEKAAWSAzureGoogle CloudOracle Cloud
Industry Keywords
Distributed EnvironmentsCloud InfrastructureAI/ML PlatformsNVIDIA GPU NetworkingCollaboration Tools (Confluence, Slack)

Tech Stack

Tools & technologies
AnsibleAWSAzureCloudDNSGrafanaKubernetesLinuxNFSNode.jsOraclePrometheusPythonTCP/IP

About the role

Key responsibilities & impact
  • Run the WEKA environment day to day, including monitoring, health, upgrades, capacity, configuration, and audits
  • Create and manage file systems, object-storage tiering, snapshots, quotas, and POSIX, NFS, SMB, and S3 protocol configuration
  • Maintain a healthy, documented, and recoverable operational run state
  • Execute approved changes with scope confirmation, risk classification, validation, rollback planning, and outcome recording
  • Monitor drift, capacity pressure, and anomalies and raise issues with recommended actions
  • Work with customer tenants and platform teams to understand storage usage for training, inference, checkpointing, data loading, and scheduling
  • Tune WEKA for workloads through filesystem layout, tiering, snapshot policy, quotas, and per-tenant QoS
  • Validate and standardize GPU node clients, BlueField-3 and SR-IOV networking, mount options, cgroups, kubelet CPU policy, and WEKA Kubernetes operator and CSI driver
  • Onboard tenants and clients using repeatable, verified procedures
  • Establish performance baselines, run benchmarks and tests, and report changes and causes
  • Identify scaling limits involving client counts, process ceilings, NAT, and routing across tenant networks
  • Build and maintain customer runbooks covering procedures, access, validation, rollback, and recovery
  • Automate repeated manual work with scripts, automation, or checklists using Python, Bash, and Ansible
  • Prepare handovers and backup briefs
  • Explain procedures and decisions to customer operators
  • Open and update cases and follow escalation paths to Technical Services, GRE, and R&D
  • Keep the WEKA team informed of changes, risks, and customer communications
  • Collaborate with cloud providers, network teams, and third parties to resolve issues
  • Contribute operational results, trends, and risks to weekly status and quarterly reviews

Requirements

What you’ll need
  • 10+ years administering enterprise infrastructure at scale, with strong storage experience (file, object, or block)
  • Deep Linux administration skills, including RHEL-family and Ubuntu, in distributed environments
  • Hands-on experience with high-speed networking: Ethernet and InfiniBand, RDMA, DNS, ACLs, and TCP/IP fundamentals
  • Working knowledge of Kubernetes and containers, including persistent storage, CSI, and operators
  • Experience operating GPU, HPC, or large cloud infrastructure for external customers or multiple tenants
  • Experience with at least one major cloud: AWS, Azure, Google Cloud, or Oracle Cloud
  • Proficiency in Python, Bash, or Ansible
  • Experience with monitoring and alerting stacks, including Prometheus, Grafana, and log and metrics tooling
  • Experience with ticketing systems such as Jira and discipline in major incident management
  • Clear written and verbal communication
  • Strong technical writing
  • Prior experience with WEKA or another parallel or distributed file system is nice to have
  • Experience running Slurm or Kubernetes-based AI/ML platforms is nice to have
  • Experience with NVIDIA GPU cluster networking and DGX-class systems is nice to have
  • Experience collaborating between support and product or engineering teams is nice to have
  • Familiarity with Confluence, Slack, and other collaboration tools is nice to have

Benefits

Comp & perks
  • No specific benefits, perks, or compensation extras stated