Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Megaport

Senior Site Reliability Engineer

Megaport

. Improve production reliability and system resilience within an SRE-scoped team .

Posted 9/15/2026full-timeRemote • Texas • United StatesSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Linux system administration, cloud infrastructure management, and automation practices while ensuring system reliability and security. Proficient in incident response and observability, with a strong focus on collaboration and stakeholder engagement.

Highest-signal resume keywords
Linux System AdministrationCloud Infrastructure ManagementKubernetes FundamentalsInfrastructure-as-Code (Terraform)CI/CD and Version Control (GitHub)

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Linux Systems AdministrationAutomationKubernetesBash ScriptingPythonGoInfrastructure-as-CodeTerraformCI/CDDatabase Management (Postgres, Cassandra, ClickHouse)
Soft Skills
CommunicationProblem-SolvingCollaborationSelf-Directed
Tools & Technologies
AWSGitHubObservability StacksMetricsLogsTraces
Industry Keywords
SRESLIsSLOsSLAsError BudgetsBlameless PostmortemsIncident ResponseProduction Environments

Tech Stack

Tools & technologies
AWSCassandraCloudKubernetesLinuxPostgresPythonTerraformGo

About the role

Key responsibilities & impact
  • Improve production reliability and system resilience within an SRE-scoped team
  • Champion high standards and industry best practices
  • Communicate with teams and stakeholders throughout projects
  • Bring fresh ideas and encourage others
  • Investigate complex technical problems
  • Work across numerous technologies in a fast-changing industry
  • Participate in on-call rotation, incident response, and blameless post-incident reviews
  • Write code, handle alerts, improve solutions, and support others
  • Engage stakeholders in requirements analysis and demonstrations
  • Ensure systems are secure, maintainable, and available
  • Support customer success and company goals

Requirements

What you’ll need
  • 5+ years administering Linux systems and related infrastructure in production environments
  • Familiarity with SLIs, SLOs, SLAs, error budgets, blast radius, and blameless postmortems
  • Focus on automation, reducing toil, and preventing problem recurrence
  • Track record of writing runbooks for broader teams
  • Strong Kubernetes and broader ecosystem fundamentals
  • Cloud infrastructure experience; AWS strongly preferred
  • Bare-metal experience is a bonus
  • Strong tool development using Bash, plus Python or Go preferred, or similar
  • Infrastructure-as-code tooling experience; Terraform preferred
  • CI/CD and version control experience; GitHub preferred
  • Database experience with Postgres, Cassandra, or ClickHouse preferred
  • Experience operating production observability stacks covering metrics, logs, and traces
  • Strong troubleshooting instincts and ownership of incident response
  • History of continual professional development
  • Self-directed style suited to an async, globally distributed team
  • Comfortable picking up adjacent work when needed

Benefits

Comp & perks
  • Flexible working environment – a remote-first culture with coworking options available.
  • 4 weeks of paid annual leave
  • Parental leave
  • Birthday leave
  • Purchased annual leave program
  • Wellness allowance
  • Employee wellbeing initiatives
  • Generous study and training allowance
  • 5 days of paid study leave
  • Creative, modern workspaces
  • Recognition programs, including Legend and Kudos awards