Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Uvation

AI Solution Architect

Uvation

. Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness.

Posted 9/18/2026full-timeRemote • IndiaSeniorLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in AI/ML architecture, including GPU compute and networking design, while ensuring high-performance and secure enterprise solutions. Proficient in developing reference architectures and leading cross-functional collaboration for deployment and operational readiness.

Highest-signal resume keywords
AI/ML Architecture ExpertiseGPU Compute Architecture DesignKubernetes and Container RuntimesNVIDIA CertificationsCloud Solutions Architecture

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
AI Factory ArchitectureGPU Scheduling and Multi-TenancyHigh-Performance NetworkingParallel File SystemsData Lake ConceptsModel Serving and Inference PlatformsPerformance EngineeringDisaster Recovery PlanningCapacity PlanningInfrastructure Telemetry
Soft Skills
Technical LeadershipCollaborationProblem-SolvingCommunication
Tools & Technologies
NVIDIA HGX/DGXCephWEKAPrometheusGrafanaOpenTelemetryAWS AI InfrastructureAzure AI InfrastructureInfiniBandNVIDIA ConnectX
Certifications & Qualifications
AWS Solutions ArchitectAzure Solutions ArchitectTOGAFCCNP/CCIECISSPKubernetes CertificationsRed Hat/Linux Certifications
Industry Keywords
AI SolutionsEnterprise ArchitectureHigh-Performance ComputingCloud SecurityZero Trust SecurityData ProtectionOperational ResilienceDisaster RecoveryCapacity ModelsPerformance Objectives

Tech Stack

Tools & technologies
AWSAzureCloudGrafanaKubernetesLinuxNFSPrometheusPyTorchTensorflow

About the role

Key responsibilities & impact
  • Own end-to-end architecture for AI Factory and enterprise AI solutions from requirements through production readiness.
  • Assess AI/ML workload requirements for training, fine-tuning, inference, batch processing and high-performance computing.
  • Design GPU compute architectures including NVIDIA HGX/DGX/OEM platforms, multi-GPU systems, NVLink/NVSwitch and GPU resource allocation.
  • Design high-performance AI networking using 100/200/400/800G Ethernet, EVPN/VXLAN and leaf-spine architectures.
  • Design AI storage and data architectures using object storage and parallel file systems such as Ceph and WEKA.
  • Define AI platform architecture across Kubernetes, HPC, container runtimes, model-serving platforms and enterprise AI frameworks.
  • Establish architecture standards for security, identity, tenant isolation, data protection, observability, disaster recovery and operational resilience.
  • Develop reference architectures, high-level/low-level designs, capacity models, bills of materials, technology evaluations and implementation roadmaps.
  • Lead technical evaluations, proof-of-concepts, vendor assessments and architecture review boards.
  • Collaborate with infrastructure, network, security, storage, cloud, data, application and operations teams.
  • Define performance, availability, scalability, security and cost objectives and validate architecture against measurable acceptance criteria.
  • Provide technical leadership during deployment, migration, integration, troubleshooting and production transition.
  • Produce AI Factory reference architectures, solution blueprints, HLDs, LLDs, architecture diagrams, capacity/performance/scalability models, technology evaluations, security architecture inputs, BOMs, sizing, migration strategies, implementation roadmaps, operational readiness checklists, runbooks and acceptance criteria.

Requirements

What you’ll need
  • 10+ years of infrastructure, cloud, enterprise architecture or solution architecture experience, with significant AI/GPU infrastructure exposure.
  • Proven experience designing large-scale enterprise platforms and translating business requirements into technical architectures.
  • Hands-on understanding of physical infrastructure, GPU systems, networking, storage and Linux platforms.
  • Bachelor's degree in Computer Science, Engineering, Information Technology or related field preferred.
  • NVIDIA certifications or equivalent GPU/AI infrastructure credentials preferred.
  • AWS Solutions Architect / Azure Solutions Architect certification preferred.
  • TOGAF or equivalent enterprise architecture certification preferred.
  • CCNP/CCIE or equivalent networking certification preferred.
  • CISSP or equivalent security certification preferred.
  • Kubernetes certifications such as CKA/CKAD preferred.
  • Red Hat / Linux certifications preferred.
  • AI/ML architecture expertise, including NVIDIA AI Enterprise, NGC, CUDA, NCCL, DCGM, GPU Operator and AI platform ecosystems.
  • PyTorch, TensorFlow and JAX, with operational understanding of training and inference workloads.
  • GPU scheduling, multi-tenancy, MIG/vGPU, GPU utilization and workload placement.
  • LLM, generative AI, RAG, fine-tuning, model serving and inference architecture.
  • NVIDIA A100/H100/H200/B200 or equivalent GPU platforms; NVLink, NVSwitch, PCIe topology and multi-GPU performance architecture.
  • DGX/HGX/OEM GPU server architecture and lifecycle management.
  • AI Factory capacity planning, rack density, power, cooling, commissioning and lifecycle strategy.
  • 100/200/400/800G Ethernet, InfiniBand, RoCEv2, RDMA, Netris, BGP, EVPN/VXLAN, VRF, ECMP, VLAN, MTU, PFC, ECN and QoS.
  • NVIDIA ConnectX/SuperNIC, Spectrum/Spectrum-X, Quantum and BlueField DPU technologies.
  • GPU east-west traffic, GPUDirect RDMA and network performance troubleshooting.
  • Parallel file systems, object storage, NFS, NVMe/NVMe-oF and high-throughput data pipelines.
  • Ceph, WEKA, VAST, Dell PowerScale, Pure FlashBlade, NetApp or equivalent technologies.
  • Data lake/lakehouse concepts, metadata, lineage, data movement and data lifecycle.
  • GPUDirect Storage and storage/network performance optimization.
  • Kubernetes, GPU Operator, container runtimes, Kubernetes GPU scheduling and HPC or equivalent workload schedulers.
  • Model serving/inference platforms, MLOps platform architecture, API gateways, service discovery, secrets management and platform integration.
  • AWS and/or Azure AI infrastructure and security services.
  • Hybrid cloud connectivity, IAM, private networking, cloud storage, workload placement, cloud cost optimization, capacity planning and FinOps.
  • Zero Trust, network segmentation, IAM/RBAC, PAM, workload identity, GPU/DPU/container/Kubernetes/firmware/supply-chain security.
  • Encryption at rest/in transit, secrets management, audit logging, compliance controls, data/model protection, tenant isolation and secure model access.
  • Prometheus, Grafana, OpenTelemetry, NVIDIA DCGM and infrastructure telemetry.
  • Monitoring across GPU, CPU, memory, network, storage, power and thermal domains.
  • High availability, backup/restore, disaster recovery, business continuity, failure-domain design, performance engineering, bottleneck analysis, SLO/SLA design and capacity forecasting.