FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesAWSCloudGrafanaKubernetesLinuxPostgresPrometheusPythonTerraform
About the role
Key responsibilities & impact- Monitor key SaaS application metrics, logs, and alerts to proactively identify and prevent service disruptions
- Support a 24/7 operational model to ensure high availability and performance of the SaaS environment
- Serve as a primary point of contact for technical customer inquiries and issues
- Partner with customers to understand requirements, troubleshoot effectively, and provide clear, concise communication
- Triage, escalate, and own technical incidents through resolution
- Analyze metrics, logs, and incident reports to provide actionable insights to engineering teams
- Identify recurring patterns and opportunities for automation and continuous improvement
- Partner cross-functionally to improve platform reliability and performance
- Design and implement automation to streamline incident response and improve system reliability
- Introduce and mature SRE best practices, including enhanced monitoring, alerting, and self-healing capabilities
- Contribute to building scalable, resilient infrastructure to support AI inference workloads
Requirements
What you’ll need- Bachelor’s degree in Computer Science, Information Technology, or a related field (or equivalent practical experience)
- 1–3+ years of experience in technical support, systems administration, or a similar role
- Strong understanding of SaaS environments and cloud-based architectures, preferably AWS
- Proficiency in at least one scripting language, e.g., Python
- Solid understanding of web technologies, including HTTP, REST APIs, and JSON
- Experience working with ticketing systems
- Strong problem-solving and analytical skills
- Excellent written and verbal communication skills
- Ability to work independently and collaboratively in a team environment
- Willingness to learn new technologies and adapt to evolving requirements
- Experience with monitoring and observability tools, e.g., Prometheus and Grafana
- Familiarity with configuration management tools, e.g., Terraform
- Experience working with cloud infrastructure technologies
- Exposure to SRE principles and reliability engineering practices
- Strong understanding of networking fundamentals
- Experience with databases, PostgreSQL, operating systems, Linux, and Kubernetes
Benefits
Comp & perks- Incentive compensation
- Bonus
- Restricted stock units
- Benefits package
- Reasonable accommodations for candidates
