FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesGrafanaKafkaMySQLPHPPostgresPrometheusPuppetPythonGo
About the role
Key responsibilities & impact- Own and evolve comprehensive monitoring and alerting for MySQL InnoDB Clusters and PostgreSQL infrastructure across multiple datacenters
- Establish and execute a quarterly backup verification and disaster recovery testing program across all database systems
- Serve as first-tier on-call responder for database incidents, executing documented runbooks and escalating to senior engineering when architecture-level decisions are required
- Monitor and maintain replication health across databases
- Own database user lifecycle management, including provisioning, deprovisioning, access audits, and role-based access control
- Execute and track security compliance remediation, including pen-test findings, GRC audit requirements, and encryption-at-rest verification
- Manage operational health of the Debezium/Kafka Connect data pipeline in coordination with the Kafka infrastructure team
- Build and maintain Puppet profiles for database infrastructure configuration management
- Write PHP, Python, or Go automation tooling to reduce operational toil
- Develop and maintain runbooks, operational documentation, and disaster recovery procedures
- Partner with the Senior Platform Engineer (Databases) as a peer, reviewing work, sharing on-call, and splitting ownership of database reliability across production systems
Requirements
What you’ll need- 7+ years of experience in Site Reliability Engineering, DevOps, or Database Operations roles in production environments at scale
- Deep operational expertise with MySQL in production — InnoDB Cluster, Group Replication, MySQL Router, and ProxySQL — with strong troubleshooting skills for replication, performance, and reliability issues
- Production experience with PostgreSQL — replication, high availability, performance tuning, and operational management
- Strong proficiency with configuration management tools (Puppet preferred) and infrastructure-as-code practices
- Experience with database backup tools (xtrabackup, pg_dump, mysqldump) and disaster recovery procedures
- Proficiency in PHP and Python or Go for automation, tooling, and integration with existing codebases
- Experience with database observability — Prometheus exporters, Grafana dashboards, alerting frameworks, and SLO/error-budget methodology
- Familiarity with Kafka Connect, Debezium, or similar change-data-capture pipelines
- Strong incident response skills with experience in on-call rotations, including post-incident review and remediation
- Excellent communication skills and ability to collaborate across engineering teams as a senior peer
- Must be residing in one of the listed U.S. states
- Must be legally authorized to work in the United States
Benefits
Comp & perks- 100% company-paid insurance premiums for employee medical, dental and vision plans
- 401(k) plan that matches 100% up to 4%, with immediate vesting
- Professional Development Reimbursement of $2,500 each year
- 11 Holidays + Paid Time Off Accrual + Rollover Plan
- Increased PTO at 3 year and 10 year anniversary
- 1 month paid sabbatical every 5 years
- Anniversary Bonus each year
- $500 stipend for remote office setup in first year + $400 each following year
- Internet reimbursement up to $75 per month
- Gym membership reimbursement up to $50 per month
- Company paid Wellable subscription
