FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing, deploying, and maintaining cloud and on-premises infrastructure for big data and Generative AI workloads, with a strong focus on security, scalability, and observability. Proficient in using tools like Terraform, Ansible, and Kubernetes to optimize resource utilization and manage CI/CD pipelines effectively.
Highest-signal resume keywords
AWS Cloud Infrastructure ManagementKubernetes Application DeploymentTerraform and Ansible ProficiencyCI/CD Pipeline Development with GitLab CIObservability Implementation with Prometheus
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Linux System AdministrationBig Data PlatformsGenerative AI WorkloadsNetworking and SecurityPerformance TuningMLOps PracticesHadoop Data ProcessingSpark Data ProcessingSnowflake AdministrationAI Tools Utilization
Soft Skills
Analytical Problem-SolvingCollaboration with Security TeamsCuriosity about Generative AI
Tools & Technologies
AWS Secrets ManagerVaultKMSGitLab CIArgoCDPrometheusZabbixGrafanaEKSAKS
Certifications & Qualifications
AWS CertificationAzure CertificationKubernetes Certification
Industry Keywords
Cloud-Native ServicesHybrid EnvironmentsRegulated IndustriesAI Agent FrameworksDistributed Systems
Tech Stack
Tools & technologiesAnsibleAWSAzureCloudDistributed SystemsGrafanaHadoopKubernetesLinuxPrometheusSparkTerraformVault
About the role
Key responsibilities & impact- Design, deploy, and maintain cloud and on-premises infrastructure for big data platforms and Generative AI workloads
- Build secure, scalable AWS or Azure environments, including multi-account/landing-zone setups, networking, IAM, and identity integration
- Deploy, scale, and operate production LLM-powered services, including agent frameworks, RAG pipelines, vector stores, and orchestration layers
- Ensure AI services are reliable, observable, and cost-efficient
- Use Terraform and Ansible for repeatable deployments
- Manage secrets, keys, and AI provider credentials securely using AWS Secrets Manager, Vault, and KMS
- Build and maintain CI/CD pipelines with GitLab CI and ArgoCD for infrastructure, data applications, and AI services
- Manage automated testing, static analysis, AI-assisted code review, and release processes
- Deploy and operate EKS/AKS Kubernetes clusters for big data and AI workloads
- Optimize GPU/CPU resource utilization for inference and batch jobs
- Implement and optimize observability using Prometheus, Zabbix, Grafana, and cloud-native tools
- Monitor LLM-specific signals including model latency, error rates, prompt/response quality, and cost
- Collaborate with network and security teams on secure connectivity, VPNs, Transit Gateway, and private endpoints
- Apply compliance best practices in regulated industries
- Identify and resolve performance bottlenecks across big data clusters and AI inference workloads
- Use AI tools such as Claude Code and MCP-based assistants for code review, IaC generation, runbook automation, incident response, and team adoption
Requirements
What you’ll need- 3+ years of hands-on experience in DevOps or SRE roles
- Experience ideally on big data, distributed systems, or modern AI/ML platforms
- Strong knowledge of Linux (RHEL), including scripting, system administration, and troubleshooting
- Hands-on experience with AWS or Azure, including deployment of cloud-native services and infrastructure
- Expertise deploying, managing, and scaling applications using Kubernetes
- Proficiency with Terraform and Ansible
- Experience with CI/CD pipelines and GitLab CI, ArgoCD, or similar
- Proficiency with Prometheus or Zabbix for metrics and alerting
- Strong understanding of networking, security, VPNs, and performance tuning in hybrid environments
- Strong analytical and problem-solving skills, with experience resolving complex platform and performance issues
- Curiosity about Generative AI and willingness to apply AI tools, coding agents, and MCP to engineering work
- Experience deploying or integrating MCP servers or AI agent frameworks is nice to have
- MLOps practices are nice to have
- Experience managing Cloudera or similar on-premises big data platforms is nice to have
- Experience with Hadoop, Spark, or similar data processing frameworks is nice to have
- Snowflake administration and automation is nice to have
- AWS, Azure, or Kubernetes certifications are nice to have
- Security best practices and tools for cloud and on-premises environments are nice to have
Benefits
Comp & perks- Participation in the company's stock options program
- Flexible Benefits & Personal learning budget
- 10 Growth Days per year - dedicated time for learning and development
- Hybrid work environment
- All the support you need from our experienced team to become an even better professional
- Direct involvement in building infrastructure behind real, production Generative AI products
- Daily access to and active use of modern AI tooling as part of your engineering toolkit
- Ownership and dynamics in your role
- International team
