Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
OutSystems

Senior Site Reliability Engineer

OutSystems

. Lead and onboard services and teams to reliability tenets .

Posted 9/17/2026full-timeLisbon • PortugalSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering with a focus on designing and implementing scalable cloud-native infrastructure, managing SLAs and SLOs, and automating operational tasks. Proficient in Python programming and experienced in monitoring and troubleshooting complex distributed systems.

Highest-signal resume keywords
Site Reliability EngineeringCloud-Native Infrastructure DesignPython ProgrammingKubernetes ManagementMonitoring and Troubleshooting

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Site Reliability EngineeringCloud-Native InfrastructurePythonKubernetesLinuxNetworkingContainersAutomationInfrastructure as CodeDebugging
Soft Skills
CommunicationCollaborationContinuous ImprovementTroubleshootingKnowledge Sharing
Tools & Technologies
HadoopAWSGrafanaELK StackPrometheusAI Native IDEsGitHub CoPilotCursorClaudeBash/Shell Scripting
Certifications & Qualifications
BS/MS in Computer Science
Industry Keywords
Service Level ObjectivesService Level AgreementsIncident ResponseRoot-Cause AnalysisPost-MortemsSRE KPIsMTTAMTTRSLO CoverageDetection Ratio

Tech Stack

Tools & technologies
AWSCloudDistributed SystemsGrafanaHadoopKubernetesLinuxPrometheusPythonShell ScriptingGo

About the role

Key responsibilities & impact
  • Lead and onboard services and teams to reliability tenets
  • Establish and maintain Service Level Objectives and Service Level Agreements
  • Design and implement scalable, reliable, and secure cloud-native infrastructure
  • Collaborate with software development teams to build observable, fault-tolerant, recoverable, scalable, and performant systems
  • Implement monitoring, alerting, logging, and tracing solutions
  • Lead incident response, resolve incidents quickly, and conduct root-cause analyses and post-mortems
  • Automate operational tasks, focusing on rapid incident detection and recovery
  • Program in Python with generative AI tooling to develop mission-critical automation and tools
  • Foster continuous improvement and knowledge sharing
  • Communicate system reliability and performance updates to stakeholders
  • Participate in an on-call rotation providing 24/7 production support
  • Track SRE KPIs including SLA/SLO compliance, SLO coverage and detection ratio, MTTA, and MTTR

Requirements

What you’ll need
  • BS/MS in Computer Science or equivalent
  • 6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale
  • History of end-to-end project delivery
  • Experience managing Hadoop and Kubernetes infrastructure and related services, or equivalent experience
  • Advanced knowledge of Linux, Networking, and Containers
  • Proficiency in at least one high-level programming language, such as Python or GoLang
  • Strong troubleshooting and debugging skills
  • Fluency in English
  • Understanding or hands-on experience with prompt engineering in software development
  • Familiarity with AI Native IDEs or AI Assistants such as Cursor, GitHub CoPilot, and Claude
  • Ability to establish, monitor, and improve SLOs, SLIs, and SLAs
  • Experience with containerization and orchestration platforms, mainly Kubernetes and EKS
  • Experience with automation and Infrastructure as Code tools
  • Experience with Python, Go, Bash/Shell scripting, or other automation tools/languages
  • Familiarity with AWS services
  • Proficiency in monitoring and troubleshooting complex distributed systems
  • Experience with Grafana, ELK stack, Prometheus, or similar tools
  • Strong understanding of resilient and fault-tolerant system design
  • Expertise in debugging complex distributed systems

Benefits

Comp & perks
  • Professional Development Fund
  • Internal Mobility Program
  • Structured professional development programs
  • Inclusive culture
  • Global collaboration with accessible mentors
  • Equal opportunity employment