FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in managing production system reliability, scalability, and observability, with a strong focus on AWS infrastructure, Kubernetes, and Terraform. Proficient in automating deployment pipelines and ensuring compliance and security best practices in a fast-paced environment.
Highest-signal resume keywords
AWS Infrastructure ManagementKubernetes DeploymentTerraform Infrastructure as CodeObservability with PrometheusIncident Response Leadership
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
SREDevOpsInfrastructure EngineeringPython ScriptingBash ScriptingCI/CD PipelinesGitOpsDatabase ReliabilityReal-Time Data ProcessingNetworking
Soft Skills
CollaborationProblem-SolvingAdaptability
Tools & Technologies
PrometheusGrafanaDatadogGitHub ActionsArgoCDFlux
Industry Keywords
EnergyClimate TechCritical Infrastructure
Tech Stack
Tools & technologiesAWSCloudDistributed SystemsFluxGrafanaKubernetesPrometheusPythonTerraform
About the role
Key responsibilities & impact- Own the reliability, scalability, and observability of production systems
- Design and operate infrastructure on AWS using Terraform and Kubernetes
- Build monitoring, alerting, and observability with Prometheus, Grafana, Datadog, or similar tools
- Define meaningful SLOs and SLIs
- Automate deployment pipelines, capacity management, and self-healing systems
- Partner with engineering on architecture reviews to identify reliability and scalability risks
- Manage database and data pipeline reliability for large-scale, real-time grid data processing
- Drive security and compliance best practices across infrastructure
- Run on-call for production systems and lead incident response
Requirements
What you’ll need- 5+ years in SRE, DevOps, or infrastructure engineering roles
- Deep experience with Kubernetes, Terraform/IaC, and cloud platforms; AWS preferred
- Strong scripting/programming ability in Python and Bash
- Observability experience with Prometheus, Grafana, and Datadog
- Track record of running on-call for production systems and leading incident response
- Experience with CI/CD pipelines, including GitHub Actions, and infrastructure automation
- Experience with GitOps concepts and tooling such as ArgoCD or Flux
- Solid understanding of networking, distributed systems, and database reliability
- Comfortable operating in a fast-moving startup environment with ambiguity
- Experience with data-intensive or real-time processing systems preferred
- Background in energy, climate tech, or critical infrastructure preferred
- Experience scaling infrastructure through hypergrowth preferred
- On-prem Kubernetes deployment experience preferred
- Windows Server administration experience preferred
Benefits
Comp & perks- Performance bonus
- Equity
- Comprehensive health, dental, and vision coverage
- Lunch provided three days a week in office
- Hybrid schedule for local employees: 3 days in office and 2 days remote
- Access to leading academic, industry, and government partners in the AI-energy ecosystem
- Mission-driven team focused on shaping the future of the energy transition
