FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering (SRE) with a focus on reliability improvements, incident management, and operational excellence. Proficient in leveraging cloud technologies, particularly Microsoft Azure, to enhance system performance and resilience.
Highest-signal resume keywords
Site Reliability Engineering (SRE)Microsoft AzureIncident ManagementAI-Assisted DevelopmentInfrastructure-as-Code
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Reliability EngineeringToil-Reduction ProjectsFull-Stack ObservabilityPerformance OptimizationDisaster Recovery PlanningTerraformAzure DevOps PipelinesGitHubAzure MonitorLog Analytics
Soft Skills
MentoringCollaborationCoaching
Tools & Technologies
Azure VMs/VMSSAzure FunctionsAzure Kubernetes Service (AKS)DatadogAI/ML Tools
Certifications & Qualifications
Public Trust Security Clearance
Industry Keywords
Incident AnalysisBlameless Post-MortemsRoot-Cause AnalysisHigh-Availability SystemsResiliency Patterns
Tech Stack
Tools & technologiesAzureKubernetesSDLCTerraform
About the role
Key responsibilities & impact- Drive day-to-day operations for the Life Sciences SRE team
- Support priorities set by the SRE Manager
- Deliver reliability and toil-reduction initiatives with SREs and contractors
- Partner with the Senior Software Architect for Life Sciences to modernize processes, tooling, and development strategy
- Collaborate with SREs across Reed Tech to converge on common tools, standards, and practices
- Influence service level objectives across Life Sciences production systems
- Participate in major incident resolution as an escalation point and guide restoration efforts
- Troubleshoot and resolve complex systems and application issues across development and production environments
- Improve the SRE framework and contribute to shared SRE knowledge documentation
- Create disaster recovery plans
- Mentor and prepare junior SREs for on-call readiness
- Champion AI-assisted development and operational tooling
- Implement reusable dashboards, observability standards, SLOs, and error budgets
- Conduct incident analysis, performance optimization, blameless post-mortems, and root-cause analysis
- Support platform modernization, CI/CD and SDLC improvements, standard platform adoption, and operational health reporting
- Advise on reliability improvements and SRE standards
- Promote SRE best practices, coach engineers, identify skill gaps, and champion automation and toil reduction
Requirements
What you’ll need- U.S. citizenship required
- Must successfully pass a background investigation and achieve Public Trust security clearance
- Preference for candidates located near the Horsham, PA office for a hybrid onsite schedule
- Experience as an SRE and ability to serve as a technical resource for reliability engineering
- Experience owning complex reliability and toil-reduction projects
- Experience acting as an escalation point during major incidents
- Experience with Microsoft Azure
- Experience with Azure VMs/VMSS, App Service, Azure Functions, and growing use of Azure Kubernetes Service (AKS)
- Experience with Azure DevOps Pipelines
- Experience with source control and GitHub; GitHub Actions under evaluation
- Experience with Azure Monitor and Log Analytics; Datadog supplemental
- Experience with Terraform
- Experience using AI-assisted coding tools such as Claude, GitHub Copilot, or Codex
- Interest or hands-on experience applying AI/ML to observability, automation, or incident response
- Comfort experimenting with and evaluating emerging AI tooling
- Strong understanding of full-stack observability, incident analysis, and performance optimization
- Experience with on-call readiness, mentoring, major incident response, blameless post-mortems, and root-cause analysis
- Advanced knowledge of high-availability systems, resiliency patterns, deployment strategies, and recovery practices
- Experience with failover testing, production recovery, and automating recovery processes using Infrastructure-as-Code and configuration management tools
Benefits
Comp & perks- Comprehensive, multi-carrier program for medical, dental and vision benefits
- 401(k) with match
- Employee Share Purchase Plan
- Wellness platform with incentives
- Headspace app subscription
- Employee Assistance and Time-off Programs
- Short-and-Long Term Disability, Life and Accidental Death Insurance, Critical Illness, and Hospital Indemnity
- Family Benefits, including bonding and family care leaves, adoption and surrogacy benefits
- Health Savings, Health Care, Dependent Care and Commuter Spending Accounts
- Up to two days of paid leave each to participate in Employee Resource Groups and to volunteer with your charity of choice
- Flexible hours
- Shared parental leave
- Study assistance
- Sabbaticals
