Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Propio Aruba Realty

AI Infrastructure Engineer

Propio Aruba Realty

. Design, deploy, and operate low-latency streaming inference systems on AWS for LLM, ASR, TTS, and multimodal models .

Posted 9/29/2026full-timeOverland Park • Kansas • United StatesMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and operating low-latency streaming inference systems on AWS, with a strong focus on AI/ML infrastructure, model optimization, and production readiness. Proficient in implementing secure and efficient deployment workflows, incident response, and performance benchmarking.

Highest-signal resume keywords
AWS Infrastructure ManagementLow-Latency Streaming InferenceModel Optimization and DeploymentDistributed Systems ExperienceIncident Response and Capacity Planning

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
AI/ML InfrastructureInference PlatformsModel ServingQuantizationReal-Time Audio PipelinesGPU WorkloadsPerformance BenchmarkingDistributed TrainingStreaming ProtocolsVersioning and Lineage
Soft Skills
OwnershipCollaborationInfluencing DecisionsDriving Business Results
Tools & Technologies
AWS EKSAWS EC2AWS S3AWS IAM/KMSCloudWatchOpenTelemetryVLLMTensorRT-LLMTritonONNX Runtime
Industry Keywords
Low-Latency InferenceReal-Time FactorStreaming InferenceSecure AI PlatformPHI/PII ComplianceFSDPDeepSpeedMegatron-LMSlurmEdge AI Deployment

Tech Stack

Tools & technologies
AWSDistributed SystemsEC2GRPCRay

About the role

Key responsibilities & impact
  • Design, deploy, and operate low-latency streaming inference systems on AWS for LLM, ASR, TTS, and multimodal models
  • Optimize time to first token/audio, p95/p99 end-to-end latency, real-time factor, throughput, concurrency, GPU utilization, and cost per stream
  • Support and evolve the research team’s AWS-based training environment
  • Provide reproducible containers, GPU job scheduling, distributed job execution, checkpoint and resume capabilities, experiment tracking, model and data artifact access, and researcher self-service
  • Build model and artifact registries, lineage and versioning, automated evaluation gates, deployment pipelines, shadow and canary releases, rollback workflows, runtime and configuration management, and production observability
  • Partner with researchers and device and embedded teams to build model optimization, packaging, validation, and deployment workflows for resource-constrained edge targets
  • Support model export and compilation, post-training quantization, and runtime integration
  • Own capacity planning, production readiness, incident response, disaster recovery, and secure AI platform operation
  • Apply least-privilege IAM, KMS encryption, private networking, and secrets management
  • Take ownership of important initiatives and outcomes, drive business results, influence decisions, and partner with high-performing team members

Requirements

What you’ll need
  • 3+ years of experience building and operating AI/ML infrastructure, inference platforms, or distributed systems in production
  • Hands-on production experience with AWS, particularly EKS, EC2 GPU workloads, ECR, S3, IAM/KMS, VPC networking, and CloudWatch and/or OpenTelemetry
  • Hands-on experience with at least one inference stack, such as vLLM, SGLang, TensorRT-LLM, Triton, KServe, or Ray Serve
  • Experience operating ML/LLM systems in production, including model serving, autoscaling, monitoring, incident response, and performance benchmarking
  • Familiarity with GPU infrastructure and at least one serving stack, such as vLLM, SGLang, TensorRT-LLM, or Triton
  • A working understanding of LLM Ops practices, including evaluation, observability and tracing, cost control, and versioning
  • Low-latency, real-time, or streaming inference experience, especially for audio/speech—directly relevant to interpretation
  • Experience with real-time audio pipelines and streaming protocols, including WebRTC, WebSocket, and gRPC streaming
  • Experience supporting distributed training environments using FSDP, DeepSpeed, Megatron-LM, Slurm, EKS, or SageMaker HyperPod
  • Inference optimization experience, including quantization, KV-cache optimization, speculative decoding, and continuous batching
  • Experience building Edge AI deployment toolchains using ONNX Runtime, TensorRT/Jetson, ExecuTorch, llama.cpp, MLC-LLM, or similar runtimes
  • Experience with SRE practices, capacity planning, production incident response, and secure infrastructure for PHI/PII
  • Contributions to relevant open-source infrastructure, serving, observability, or edge-runtime projects

Benefits

Comp & perks
  • Opportunities to learn, develop, and expand capabilities
  • Opportunities to build a meaningful career
  • Support for making a real difference and pursuing leadership opportunities
  • Room to grow and support for career development
  • Opportunities to build expertise and expand career path