FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesAnsibleCloudDistributed SystemsGrafanaIoTKubernetesPythonTerraform
About the role
Key responsibilities & impact- Architect the future of Omnidian’s platform reliability and operational ecosystem
- Formulate a multi-year technical roadmap for platform reliability, infrastructure, and operational excellence
- Secure executive and stakeholder buy-in for the SRE/DevOps vision and communicate progress
- Translate technical vision into actionable work plans, milestones, technical specifications, and success criteria including SLIs/SLOs and error budgets
- Design, implement, and maintain production-grade Kubernetes platforms and cluster management
- Build and evolve Infrastructure as Code and configuration management with Ansible and Terraform
- Own and improve CI/CD pipelines, deployment strategies, and release automation
- Implement and refine Grafana-centered observability stacks covering metrics, logging, and tracing
- Create automation for toil reduction, self-healing systems, and infrastructure provisioning
- Use AI coding and ops assistants for scripting, infrastructure code, refactoring, pipeline improvements, and reliability testing
- Lead agile project planning, milestone delivery, status communication, dependency management, and risk escalation
- Establish documentation practices including runbooks, architecture decision records, and post-incident reviews
- Identify and implement process, codebase, and architecture improvements
- Provide technical leadership and mentorship across teams
Requirements
What you’ll need- 14+ years of experience building, operating, and optimizing large-scale distributed systems, cloud infrastructure, and production platforms
- Deep hands-on expertise with Kubernetes, including cluster architecture, networking, scaling, security, and day-2 operations
- Strong experience with Ansible or equivalent configuration management and Terraform or equivalent Infrastructure as Code tools
- Proven ownership of modern CI/CD pipelines, deployment strategies, and release engineering practices
- Extensive experience designing and operating observability platforms, with strong proficiency in Grafana
- Solid scripting and automation skills using Python, Bash, or equivalent
- Demonstrated ability to establish and drive SRE practices including SLIs/SLOs, error budgets, incident management, and toil reduction
- Extensive use of advanced coding/ops assistants for automation, infrastructure code, testing, and establishing team standards for safe and effective use
- Experience designing systems with appropriate human oversight for automated or agentic operational workflows
- Experience integrating generative AI/LLM components or advanced automation frameworks into production infrastructure or operational systems is a plus
- Familiarity with advanced Kubernetes operators, service meshes, or multi-cluster management is a plus
- Experience implementing tracing frameworks and advanced observability for complex distributed systems is a plus
- Knowledge of event-driven architectures, time-series data, or high-cardinality metrics environments is a plus
- Exposure to IoT telemetry or similar high-volume data streams is a plus
- Solar industry experience is a plus
Benefits
Comp & perks- Monthly health insurance premiums: 100% covered for employees and 50% for dependents
- Bonuses
- Long-term stock options
- Up to $500 in annual learning reimbursement for courses, certifications, or conferences
- Annual merit increases
- Company-wide wellness and social Slack channels
- Mentorship and investment in employees
- Opportunities for career growth
