FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and automating monitoring and observability solutions using tools like Datadog, Prometheus, and Grafana, while optimizing cloud infrastructure and ensuring reliability and scalability. Proficient in software engineering with a focus on Python, TypeScript, and cloud-native technologies, alongside strong troubleshooting and mentoring capabilities.
Highest-signal resume keywords
Monitoring And Observability SolutionsCloud Infrastructure OptimizationSoftware Engineering With PythonStrong Experience With TerraformCapacity Planning And Performance Tuning
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonTypeScriptGoJavaTerraformPulumiCrossplaneCI/CDMonitoringDebugging
Soft Skills
MentoringCollaborationTroubleshootingContinuous ImprovementCommunication
Tools & Technologies
DatadogPrometheusGrafanaAWSGCPAzure
Industry Keywords
Site Reliability EngineeringCloud-Native InfrastructureSLIsSLOsIncident Response
Tech Stack
Tools & technologiesAWSAzureCloudDistributed SystemsGoogle Cloud PlatformGrafanaJavaLinuxPrometheusPythonTerraformTypeScriptGo
About the role
Key responsibilities & impact- Design, maintain, and automate monitoring, alerting, and observability solutions using Datadog, Prometheus, and Grafana
- Define and monitor SLIs/SLOs, improve incident response processes, and support on-call operations
- Lead or support production incident resolution
- Conduct blameless post-mortems and drive continuous reliability improvements
- Optimise cloud infrastructure costs while maintaining performance, scalability, and reliability
- Support capacity planning, performance tuning, and resilience initiatives through monitoring and load-testing insights
- Collaborate with engineering, QA, security, and product teams
- Contribute to system design reviews focused on reliability, scalability, and operational excellence
- Share knowledge, mentor colleagues, and contribute to the Site Reliability Engineering Community of Practice
Requirements
What you’ll need- Professional software engineering experience with Python, TypeScript, Go, Java or similar
- Best practices including CI/CD, testing, and code reviews
- Hands-on experience with AWS, GCP, or Azure
- Understanding of cloud-native infrastructure, networking, storage, and distributed systems
- Strong experience with Terraform, Pulumi, or Crossplane
- Data-driven engineering using metrics, logs, SLIs/SLOs, and error budgets
- Strong troubleshooting and debugging skills across Linux, networking, and distributed systems
- Experience with capacity planning, performance tuning, and cloud cost optimisation
- Automation-first mindset
- Passion for mentoring others and promoting Site Reliability Engineering best practices
Benefits
Comp & perks- Starting salary from €7000 gross/month, depending on experience
- Career growth through cutting-edge technologies in a global environment
- Extra vacation days
- Hybrid working options
- Modern, pet-friendly office with parking and snacks
- Ability to work remotely from another EU country for up to 30 days per year
- Brand-new, high-performance equipment
- Dynamic, collaborative, and friendly working atmosphere
- Additional health insurance
- 3rd-pillar pension benefits
