FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in cloud engineering and DevOps practices, with a strong focus on incident response, monitoring, and automation. Proficient in AWS, Kubernetes, and infrastructure as code, ensuring high availability and reliability of production systems.
Highest-signal resume keywords
AWS Production ExperienceKubernetes TroubleshootingIncident ResponseInfrastructure as Code with TerraformMonitoring and Observability Tools
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Cloud EngineeringDevOpsIncident ResponseKubernetesAWSTerraformScripting in BashPythonMonitoring with PrometheusLoad Testing
Soft Skills
Calm CommunicationTeam Collaboration
Tools & Technologies
PagerDutyGrafanaCloudWatchLokiDatadogEKSRDSVPC NetworkingCI/CD PipelinesAI Tools for SRE
Certifications & Qualifications
AWS Certification
Industry Keywords
Production Web ApplicationsPerformance IssuesReliability IssuesMonitoring WorkloadsSynthetic ChecksHealth ChecksLoad TestsPost-MortemsError BudgetsCapacity Planning
Tech Stack
Tools & technologiesAWSAzureCloudDNSEC2Google Cloud PlatformGrafanaJMeterKubernetesLinuxMySQLPostgresPrometheusPythonRedisTerraform
About the role
Key responsibilities & impact- Act as first responder for production alerts and incidents across services, from triage through resolution
- Diagnose and fix issues directly in AWS and Kubernetes, including failing pods, resource exhaustion, bad deployments, networking/DNS, database, and cache problems
- Roll back, scale, reconfigure, or patch infrastructure to restore service quickly
- Escalate to development teams only when a code change is needed, providing a clear diagnosis
- Own PagerDuty setup and business-hours incident response while reducing MTTD and MTTR
- Run blameless post-mortems and drive technical follow-up work
- Automate runbooks and repetitive operational work, including AI-assisted triage, investigation, and remediation
- Build and maintain monitoring for Kubernetes workloads and services using Prometheus/Mimir, Loki, Tempo, Grafana, and OpenTelemetry
- Monitor production releases and identify regressions in latency, errors, or resource usage
- Create and maintain synthetic checks, smoke tests, health checks, and load/performance tests
- Partner with engineering teams on performance and reliability issues and define SLOs, SLIs, and error budgets
- Harden the platform through Terraform changes, Kubernetes resource tuning, autoscaling, CI/CD checks, secrets, and IAM improvements
- Maintain service documentation and architecture decisions
Requirements
What you’ll need- 3+ years in a cloud engineering, DevOps, or SRE role supporting production web applications
- Hands-on production incident response experience, including diagnosing and fixing issues directly
- Strong hands-on AWS production experience with EKS, EC2, RDS, VPC networking, IAM, and CloudWatch
- Deep production Kubernetes experience, including troubleshooting, debugging, and monitoring workloads
- Experience identifying and resolving performance and reliability issues with engineering teams
- Experience with monitoring and observability tools such as Prometheus, Grafana, Loki, Datadog, or CloudWatch
- Experience with on-call and alerting tools such as PagerDuty
- Experience building automated production reliability tests or checks, including synthetic, smoke, health, or load tests
- Willingness to participate in a future after-hours on-call rotation
- Infrastructure as code with Terraform and comfort working in CI/CD pipelines
- Solid Linux, networking, and container fundamentals
- Scripting/automation in Bash, Python, or similar
- Calm, clear communication during incidents and across teams
- Preferred: AI tools for SRE work; Azure and possibly GCP; multiple technology stacks; AWS certification; LGTM or OpenTelemetry at scale; k6, Locust, or JMeter; MySQL/PostgreSQL operations; Redis/Memcached tuning; Cloudflare; cloud cost optimization and capacity planning
Benefits
Comp & perks- Remote position
- Preferential consideration may be given to individuals within a reasonable commuting distance of one of the offices
- Equal opportunity employer
- Unique accommodations available during the interview process
- Potential criminal background check in the final interview phase
