Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Rentsync

Site Reliability Engineer

Rentsync

. Act as first responder for production alerts and incidents across services, from triage through resolution .

Posted 9/30/2026full-timeRemote • CanadaMid-LevelSenior💰 CA$80,000 - CA$110,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in cloud engineering and DevOps practices, with a strong focus on incident response, monitoring, and automation. Proficient in AWS, Kubernetes, and infrastructure as code, ensuring high availability and reliability of production systems.

Highest-signal resume keywords
AWS Production ExperienceKubernetes TroubleshootingIncident ResponseInfrastructure as Code with TerraformMonitoring and Observability Tools

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Cloud EngineeringDevOpsIncident ResponseKubernetesAWSTerraformScripting in BashPythonMonitoring with PrometheusLoad Testing
Soft Skills
Calm CommunicationTeam Collaboration
Tools & Technologies
PagerDutyGrafanaCloudWatchLokiDatadogEKSRDSVPC NetworkingCI/CD PipelinesAI Tools for SRE
Certifications & Qualifications
AWS Certification
Industry Keywords
Production Web ApplicationsPerformance IssuesReliability IssuesMonitoring WorkloadsSynthetic ChecksHealth ChecksLoad TestsPost-MortemsError BudgetsCapacity Planning

Tech Stack

Tools & technologies
AWSAzureCloudDNSEC2Google Cloud PlatformGrafanaJMeterKubernetesLinuxMySQLPostgresPrometheusPythonRedisTerraform

About the role

Key responsibilities & impact
  • Act as first responder for production alerts and incidents across services, from triage through resolution
  • Diagnose and fix issues directly in AWS and Kubernetes, including failing pods, resource exhaustion, bad deployments, networking/DNS, database, and cache problems
  • Roll back, scale, reconfigure, or patch infrastructure to restore service quickly
  • Escalate to development teams only when a code change is needed, providing a clear diagnosis
  • Own PagerDuty setup and business-hours incident response while reducing MTTD and MTTR
  • Run blameless post-mortems and drive technical follow-up work
  • Automate runbooks and repetitive operational work, including AI-assisted triage, investigation, and remediation
  • Build and maintain monitoring for Kubernetes workloads and services using Prometheus/Mimir, Loki, Tempo, Grafana, and OpenTelemetry
  • Monitor production releases and identify regressions in latency, errors, or resource usage
  • Create and maintain synthetic checks, smoke tests, health checks, and load/performance tests
  • Partner with engineering teams on performance and reliability issues and define SLOs, SLIs, and error budgets
  • Harden the platform through Terraform changes, Kubernetes resource tuning, autoscaling, CI/CD checks, secrets, and IAM improvements
  • Maintain service documentation and architecture decisions

Requirements

What you’ll need
  • 3+ years in a cloud engineering, DevOps, or SRE role supporting production web applications
  • Hands-on production incident response experience, including diagnosing and fixing issues directly
  • Strong hands-on AWS production experience with EKS, EC2, RDS, VPC networking, IAM, and CloudWatch
  • Deep production Kubernetes experience, including troubleshooting, debugging, and monitoring workloads
  • Experience identifying and resolving performance and reliability issues with engineering teams
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Loki, Datadog, or CloudWatch
  • Experience with on-call and alerting tools such as PagerDuty
  • Experience building automated production reliability tests or checks, including synthetic, smoke, health, or load tests
  • Willingness to participate in a future after-hours on-call rotation
  • Infrastructure as code with Terraform and comfort working in CI/CD pipelines
  • Solid Linux, networking, and container fundamentals
  • Scripting/automation in Bash, Python, or similar
  • Calm, clear communication during incidents and across teams
  • Preferred: AI tools for SRE work; Azure and possibly GCP; multiple technology stacks; AWS certification; LGTM or OpenTelemetry at scale; k6, Locust, or JMeter; MySQL/PostgreSQL operations; Redis/Memcached tuning; Cloudflare; cloud cost optimization and capacity planning

Benefits

Comp & perks
  • Remote position
  • Preferential consideration may be given to individuals within a reasonable commuting distance of one of the offices
  • Equal opportunity employer
  • Unique accommodations available during the interview process
  • Potential criminal background check in the final interview phase