FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

AI Infrastructure Engineer
Propio Aruba Realty. Design, deploy, and operate low-latency streaming inference systems on AWS for LLM, ASR, TTS, and multimodal models .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and operating low-latency streaming inference systems on AWS, with a strong focus on AI/ML infrastructure, model optimization, and production readiness. Proficient in implementing secure and efficient deployment workflows, incident response, and performance benchmarking.
Highest-signal resume keywords
AWS Infrastructure ManagementLow-Latency Streaming InferenceModel Optimization and DeploymentDistributed Systems ExperienceIncident Response and Capacity Planning
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AI/ML InfrastructureInference PlatformsModel ServingQuantizationReal-Time Audio PipelinesGPU WorkloadsPerformance BenchmarkingDistributed TrainingStreaming ProtocolsVersioning and Lineage
Soft Skills
OwnershipCollaborationInfluencing DecisionsDriving Business Results
Tools & Technologies
AWS EKSAWS EC2AWS S3AWS IAM/KMSCloudWatchOpenTelemetryVLLMTensorRT-LLMTritonONNX Runtime
Industry Keywords
Low-Latency InferenceReal-Time FactorStreaming InferenceSecure AI PlatformPHI/PII ComplianceFSDPDeepSpeedMegatron-LMSlurmEdge AI Deployment
Tech Stack
Tools & technologiesAWSDistributed SystemsEC2GRPCRay
About the role
Key responsibilities & impact- Design, deploy, and operate low-latency streaming inference systems on AWS for LLM, ASR, TTS, and multimodal models
- Optimize time to first token/audio, p95/p99 end-to-end latency, real-time factor, throughput, concurrency, GPU utilization, and cost per stream
- Support and evolve the research team’s AWS-based training environment
- Provide reproducible containers, GPU job scheduling, distributed job execution, checkpoint and resume capabilities, experiment tracking, model and data artifact access, and researcher self-service
- Build model and artifact registries, lineage and versioning, automated evaluation gates, deployment pipelines, shadow and canary releases, rollback workflows, runtime and configuration management, and production observability
- Partner with researchers and device and embedded teams to build model optimization, packaging, validation, and deployment workflows for resource-constrained edge targets
- Support model export and compilation, post-training quantization, and runtime integration
- Own capacity planning, production readiness, incident response, disaster recovery, and secure AI platform operation
- Apply least-privilege IAM, KMS encryption, private networking, and secrets management
- Take ownership of important initiatives and outcomes, drive business results, influence decisions, and partner with high-performing team members
Requirements
What you’ll need- 3+ years of experience building and operating AI/ML infrastructure, inference platforms, or distributed systems in production
- Hands-on production experience with AWS, particularly EKS, EC2 GPU workloads, ECR, S3, IAM/KMS, VPC networking, and CloudWatch and/or OpenTelemetry
- Hands-on experience with at least one inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton, KServe, or Ray Serve
- Experience operating ML/LLM systems in production, including model serving, autoscaling, monitoring, incident response, and performance benchmarking
- Familiarity with GPU infrastructure and at least one serving stack, such as vLLM, SGLang, TensorRT-LLM, or Triton
- A working understanding of LLM Ops practices, including evaluation, observability and tracing, cost control, and versioning
- Low-latency, real-time, or streaming inference experience, especially for audio/speech—directly relevant to interpretation
- Experience with real-time audio pipelines and streaming protocols, including WebRTC, WebSocket, and gRPC streaming
- Experience supporting distributed training environments using FSDP, DeepSpeed, Megatron-LM, Slurm, EKS, or SageMaker HyperPod
- Inference optimization experience, including quantization, KV-cache optimization, speculative decoding, and continuous batching
- Experience building Edge AI deployment toolchains using ONNX Runtime, TensorRT/Jetson, ExecuTorch, llama.cpp, MLC-LLM, or similar runtimes
- Experience with SRE practices, capacity planning, production incident response, and secure infrastructure for PHI/PII
- Contributions to relevant open-source infrastructure, serving, observability, or edge-runtime projects
Benefits
Comp & perks- Opportunities to learn, develop, and expand capabilities
- Opportunities to build a meaningful career
- Support for making a real difference and pursuing leadership opportunities
- Room to grow and support for career development
- Opportunities to build expertise and expand career path