FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Staff Site Reliability Expert
Lightspeed Commerce. Shape the approach to infrastructure reliability, developer experience and platform capabilities .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and maintaining scalable AWS infrastructure using Infrastructure as Code, with a strong focus on observability, incident management, and disaster recovery. Proven ability to mentor engineers and improve engineering practices across cross-functional teams.
Highest-signal resume keywords
AWS Infrastructure ManagementInfrastructure as Code with TerraformDocker and Kubernetes ExperienceIncident Management and Disaster RecoveryPython, Ruby, or Go Coding
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
AWSInfrastructure as CodeTerraformDockerKubernetesPythonRubyGoShell ScriptingMySQL
Soft Skills
Strong CommunicationJudgementPrioritisation
Tools & Technologies
ECSLinuxPostgreSQLRedisDynamoDB
Industry Keywords
SaaSCloud Cost OptimisationAgileContinuous DeliveryTesting
Tech Stack
Tools & technologiesAWSCloudDockerDynamoDBKubernetesLinuxMySQLPostgresPythonRedisRubyShell ScriptingTerraformGo
About the role
Key responsibilities & impact- Shape the approach to infrastructure reliability, developer experience and platform capabilities
- Design and maintain scalable, automated AWS infrastructure using Infrastructure as Code
- Build tools and platform services enabling product engineers to ship and operate independently
- Partner with engineering teams to design resilient, secure and cost-effective systems
- Improve engineering practices and delivery processes across distributed, cross-functional teams
- Champion observability, high availability, incident management and disaster recovery
- Apply systems thinking to complex problems
- Lead incident response and drive lasting improvements
- Mentor engineers and influence reliability practices
- Participate in the on-call rotation
Requirements
What you’ll need- Strong AWS experience operating highly available, scalable production systems
- Experience in a SaaS or product-led environment
- Strong Infrastructure as Code experience with Terraform and configuration management tooling
- Experience with Docker, Kubernetes, ECS and Linux systems
- Ability to code in Python, Ruby or Go, alongside complex Shell scripting
- Experience with observability, incident management, disaster recovery and security practices
- Experience with MySQL, PostgreSQL, Redis or DynamoDB
- Experience with cloud cost optimisation
- Understanding of Agile, continuous delivery and testing
- Strong communication, judgement and prioritisation skills
- Willingness to participate in an on-call rotation
- Applicants must disclose criminal convictions and undergo criminal record checks
Benefits
Comp & perks- Lightspeed share scheme
- Unlimited paid time off policy
- Work remotely from anywhere in the world for up to 60 days per year
- Flexible working policy
- Health and wellness benefit of $500 per year
- Mental health online platform and counselling & coaching services
- Paid leave and assistance for new parents
- Paid Volunteer day
- Premium cover if you sign up for health insurance with Southern Cross
- Complimentary breakfast and lunch options
- Fresh fruits, snacks, and beverages stocked in the office
- Secure, full-time parking facilities
- Subsidised public transportation covering up to 75% of commuting costs
- Exciting events hosted regularly by the Auckland Culture Club
- High autonomy and freedom to solve complex technical problems
- Flexible, hybrid workplace model
- Genuine career growth opportunities
- Opportunity to influence technical direction and grow expertise
- Dog-friendly environment