FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Reliability Engineer
hims & hers. Own reliability for Tier 1 customer journeys, including checkout, telehealth visits, and prescription fulfillment .
Tech Stack
Tools & technologiesAWSGraphQLKubernetesPostgresTerraform
About the role
Key responsibilities & impact- Own reliability for Tier 1 customer journeys, including checkout, telehealth visits, and prescription fulfillment
- Define and instrument SLOs, golden signals, and business-level monitors
- Improve the ratio of issues detected by monitors versus humans
- Lead capacity and resilience work for high-stakes events and steady-state growth
- Design load tests and analyze prior events at second-by-second granularity
- Investigate database and backend bottlenecks
- Harden caching, GraphQL, VPC capacity, and vendor rate limits
- Mature FireHydrant, Datadog, and Jira into a connected incident-response pipeline
- Automate incident and RCA ticket creation, SLO burn-rate and composite alerting, Tier 1 alert routing, and post-mortem action tracking
- Author and maintain runbooks and severity standards
- Build and operate AI agents and tooling for OER report generation, RCA drafting, stale action-item detection, monitor and runbook gap detection, and first-pass incident triage
- Debug complex cross-boundary issues spanning frontend, API, service mesh, and database layers
- Partner with Security and product teams on anomaly detection and response
- Produce tooling and metrics for bi-weekly VP-level Operational Excellence reviews and monthly cross-engineering OERs
- Drive resulting action items to closure
- Document systems, onboard teammates, and coach engineers on SLOs, blameless post-mortems, and on-call practice
Requirements
What you’ll need- 5+ years as a Software, SRE, Platform, or Infrastructure Engineer
- Track record of owning reliability outcomes for production systems that customers depend on
- Strong software engineering fundamentals
- Ability to solve reliability problems by writing code and building tooling
- Comfortable reading application code across the stack
- Hands-on depth in observability and SLO engineering, including golden signals, burn-rate alerting, and journey-level monitors
- Production experience with AWS, Kubernetes/EKS, Terraform, and PostgreSQL (RDS/Aurora)
- Experience running or maturing incident management end to end, including on-call design, escalation policies, incident command, blameless post-mortems, and action-item follow-through
- Daily practical use of AI coding and analysis tools such as Claude or Cursor
- Judgment about when to trust AI output and when to verify it
- Communication skills to explain risk, tradeoffs, and post-incident learnings to engineers and leadership
- Preferred: Experience building AI agents or LLM-backed automation for operations
- Preferred: Load-testing and performance-engineering experience at meaningful scale, such as k6
- Preferred: Familiarity with service mesh, especially Istio
- Preferred: Background in a regulated or healthcare environment
- Preferred: Experience designing vendor and partner escalation frameworks with defined severities and response SLAs
- Must be legally authorized to work in the U.S. without restriction for any employer
- Must not require immigration sponsorship by Hims & Hers to work in the U.S.
Benefits
Comp & perks- Competitive salary
- Equity compensation
- Unlimited PTO
- Company holidays
- Quarterly mental health days
- Medical, dental, and vision benefits
- Parental leave
- Employee Stock Purchase Program (ESPP)
- 401k benefits with employer matching contribution
- Offsite team retreats
- Claude Enterprise license
- Talent-first flexible/remote work approach