FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in AI/ML architecture, including GPU compute and networking design, while ensuring high-performance and secure enterprise solutions. Proficient in developing reference architectures and leading cross-functional collaboration for deployment and operational readiness.
Highest-signal resume keywords
AI/ML Architecture ExpertiseGPU Compute Architecture DesignKubernetes and Container RuntimesNVIDIA CertificationsCloud Solutions Architecture
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AI Factory ArchitectureGPU Scheduling and Multi-TenancyHigh-Performance NetworkingParallel File SystemsData Lake ConceptsModel Serving and Inference PlatformsPerformance EngineeringDisaster Recovery PlanningCapacity PlanningInfrastructure Telemetry
Soft Skills
Technical LeadershipCollaborationProblem-SolvingCommunication
Tools & Technologies
NVIDIA HGX/DGXCephWEKAPrometheusGrafanaOpenTelemetryAWS AI InfrastructureAzure AI InfrastructureInfiniBandNVIDIA ConnectX
Certifications & Qualifications
AWS Solutions ArchitectAzure Solutions ArchitectTOGAFCCNP/CCIECISSPKubernetes CertificationsRed Hat/Linux Certifications
Industry Keywords
AI SolutionsEnterprise ArchitectureHigh-Performance ComputingCloud SecurityZero Trust SecurityData ProtectionOperational ResilienceDisaster RecoveryCapacity ModelsPerformance Objectives
Tech Stack
Tools & technologiesAWSAzureCloudGrafanaKubernetesLinuxNFSPrometheusPyTorchTensorflow
About the role
Key responsibilities & impact- Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness.
- Assess AI/ML workload requirements for training, fine-tuning, inference, batch processing and high-performance computing.
- Design GPU compute architectures including NVIDIA HGX/DGX/OEM platforms, multi-GPU systems, NVLink/NVSwitch and GPU resource allocation.
- Design high-performance AI networking using 100/200/400/800G Ethernet, EVPN/VXLAN and leaf-spine architectures.
- Design AI storage and data architectures using object storage and parallel file systems such as Ceph and WEKA.
- Define AI platform architecture across Kubernetes, HPC, container runtimes, model-serving platforms and enterprise AI frameworks.
- Establish architecture standards for security, identity, tenant isolation, data protection, observability, disaster recovery and operational resilience.
- Develop reference architectures, high-level/low-level designs, capacity models, bills of materials, technology evaluations and implementation roadmaps.
- Lead technical evaluations, proof-of-concepts, vendor assessments and architecture review boards.
- Collaborate with infrastructure, network, security, storage, cloud, data, application and operations teams.
- Define performance, availability, scalability, security and cost objectives and validate architecture against measurable acceptance criteria.
- Provide technical leadership during deployment, migration, integration, troubleshooting and production transition.
- Produce AI Factory reference architectures, solution blueprints, HLDs, LLDs, architecture diagrams, capacity/performance/scalability models, technology evaluations, security architecture inputs, BOMs, sizing, migration strategies, implementation roadmaps, operational readiness checklists, runbooks and acceptance criteria.
Requirements
What you’ll need- 10+ years of infrastructure, cloud, enterprise architecture or solution architecture experience, with significant AI/GPU infrastructure exposure.
- Proven experience designing large-scale enterprise platforms and translating business requirements into technical architectures.
- Hands-on understanding of physical infrastructure, GPU systems, networking, storage and Linux platforms.
- Bachelor's degree in Computer Science, Engineering, Information Technology or related field preferred.
- NVIDIA certifications or equivalent GPU/AI infrastructure credentials preferred.
- AWS Solutions Architect / Azure Solutions Architect certification preferred.
- TOGAF or equivalent enterprise architecture certification preferred.
- CCNP/CCIE or equivalent networking certification preferred.
- CISSP or equivalent security certification preferred.
- Kubernetes certifications such as CKA/CKAD preferred.
- Red Hat / Linux certifications preferred.
- AI/ML architecture expertise, including NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator and AI platform ecosystems.
- PyTorch, TensorFlow and JAX, with operational understanding of training and inference workloads.
- GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization and workload placement.
- LLM, generative AI, RAG, fine-tuning, model serving and inference architecture.
- NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; NVLink, NVSwitch, PCIe topology and multi-GPU performance architecture.
- DGX/HGX/OEM GPU server architecture and lifecycle management.
- AI Factory capacity planning, rack density, power, cooling, commissioning and lifecycle strategy.
- 100/200/400/800G Ethernet, InfiniBand, RoCEv2, RDMA, Netris, BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN and QoS.
- NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum and BlueField DPU technologies.
- GPU east-west traffic, GPUDirect RDMA and network performance troubleshooting.
- Parallel file systems, object storage, NFS, NVMe/NVMe-oF and high-throughput data pipelines.
- Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp or equivalent technologies.
- Data lake/lakehouse concepts, metadata, lineage, data movement and data lifecycle.
- GPUDirect Storage and storage/network performance optimization.
- Kubernetes, GPU Operator, container runtimes, Kubernetes GPU scheduling and HPC or equivalent workload schedulers.
- Model serving/inference platforms, MLOps platform architecture, API gateways, service discovery, secrets management and platform integration.
- AWS and/or Azure AI infrastructure and security services.
- Hybrid cloud connectivity, IAM, private networking, cloud storage, workload placement, cloud cost optimization, capacity planning and FinOps.
- Zero Trust, network segmentation, IAM/RBAC, PAM, workload identity, GPU/DPU/container/Kubernetes/firmware/supply-chain security.
- Encryption at rest/in transit, secrets management, audit logging, compliance controls, data/model protection, tenant isolation and secure model access.
- Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM and infrastructure telemetry.
- Monitoring across GPU, CPU, memory, network, storage, power and thermal domains.
- High availability, backup/restore, disaster recovery, business continuity, failure-domain design, performance engineering, bottleneck analysis, SLO/SLA design and capacity forecasting.
