FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesAWSCloudEC2FirewallsGrafanaJenkinsKubernetesLinuxPrometheusPythonTerraform
About the role
Key responsibilities & impact- Support and maintain scalable, highly available, and secure cloud infrastructure
- Provision and manage cloud resources using Infrastructure as Code
- Implement cloud security best practices, IAM and role-based access controls, encryption, vulnerability management, and secure infrastructure configurations
- Support containerized environments and orchestration platforms
- Apply DevSecOps principles across infrastructure and deployment workflows
- Participate in disaster recovery planning, testing, and recovery activities
- Maintain and optimize CI/CD pipelines supporting application and ML model deployments
- Improve deployment reliability and support zero-downtime deployment strategies
- Automate configuration management, infrastructure provisioning, and routine operational processes
- Troubleshoot deployment and pipeline issues and prevent recurrence
- Develop scripts and automation to reduce manual work
- Design, deploy, operate, and secure infrastructure supporting AI and agentic products
- Operate and scale ML platform infrastructure, including Databricks clusters, jobs compute, ML pipelines, and Model Serving endpoints
- Manage production model-serving infrastructure, compute capacity, provisioned throughput, and autoscaling
- Monitor model drift, data quality, inference performance, and serving health
- Maintain monitoring, logging, metrics, and alerting using Prometheus, Grafana, Coralogix, and CloudWatch
- Support incident response and perform Root Cause Analysis for infrastructure and deployment issues
- Partner with Software Engineers, ML Engineers, QA, Software Engineers in Test, IT Security, and Product Support
- Respond to engineering and Product Support requests
- Maintain accurate internal technical and operational documentation
- Collaborate with international teams across multiple time zones
Requirements
What you’ll need- 3+ years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or a similar infrastructure-focused role
- 2+ years of hands-on experience with AWS services such as EC2, S3, RDS, Lambda, IAM, VPC, SQS, and API Gateway, or similar services
- 2+ years of experience with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS
- Strong experience building and maintaining CI/CD pipelines using GitLab CI/CD or Jenkins
- Hands-on experience with Infrastructure as Code using Terraform, Terragrunt, CloudFormation, or similar technologies
- Strong Linux system administration and troubleshooting skills
- Solid understanding of networking fundamentals, including routing, load balancing, and network security
- Scripting experience with Python, Bash, or similar languages
- Hands-on experience using AI tools to improve engineering workflows, automation, troubleshooting, or agentic use cases
- Experience supporting data, ML, or other compute-intensive production workloads
- Experience with Google Cloud would be valuable
- Familiarity with Helm and service mesh technologies such as Istio, Linkerd, or Traefik would be beneficial
- Experience with serverless and event-driven architectures is a plus
- Exposure to cloud and infrastructure security practices, vulnerability management, and tools such as Nessus, Prowler, Trivy, or firewalls would be valuable
- Knowledge of security standards, compliance requirements, and cloud security best practices is beneficial
- Experience with observability, log analysis, and monitoring platforms such as Coralogix, Prometheus, or Grafana is a plus
- FinOps experience would be valuable
- Experience with API gateways or API management platforms such as Kong or Apigee is beneficial
- Experience with MLOps platforms and practices, particularly Databricks, model serving, ML pipelines, and model monitoring, would be an advantage
Benefits
Comp & perks- 15 days of vacation
- Floating and company holidays
- Wellness benefits
- Paid parental leave
- Comprehensive medical insurance
- Dental insurance
- Vision insurance
- Life insurance
- Disability insurance
- 401(k) plan
- Work from home stipend
- Learning platform, training, and tools
- Flexible work model combining in-person and remote work
