FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Site Reliability Engineer
Bank of America. Partner with engineering and technology leaders to define objective reliability goals for services .
Posted 9/18/2026full-timeJersey City • New Jersey • United StatesSenior💰 $152,600 - $191,500 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in cloud infrastructure engineering and platform reliability, with a strong focus on Google Cloud Platform (GCP) and automation practices. Proficient in designing observability solutions, implementing Infrastructure as Code (IaC) with Terraform, and leading cross-functional initiatives to enhance service reliability and operational readiness.
Highest-signal resume keywords
Google Cloud Platform (GCP)Infrastructure as Code (IaC)Terraform DevelopmentObservability Solutions DesignIncident Response and Root Cause Analysis
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Cloud Infrastructure EngineeringPlatform EngineeringTerraform Modules DevelopmentCI/CD PipelinesMonitoring and Logging ToolsDevSecOps PracticesService Level Indicators (SLIs)Service Level Objectives (SLOs)Automation FrameworksError Budget Policies
Soft Skills
Problem-Solving SkillsEffective CommunicationCollaboration with Cross-Functional TeamsAnalytical SkillsMentoring and Leadership
Tools & Technologies
Log AnalyticsDynatraceResource GraphEnterprise Monitoring ToolsPolicy-as-Code
Industry Keywords
Cloud OperationsGovernance and ComplianceIncident TriageService Health ChecksAutomation Opportunities
Tech Stack
Tools & technologiesAzureCloudDistributed SystemsDNSFirewallsGoogle Cloud PlatformTerraform
About the role
Key responsibilities & impact- Partner with engineering and technology leaders to define objective reliability goals for services
- Design observability solutions through instrumentation and dashboards
- Identify root causes of complex and high-impact issues
- Partner with cross-functional teams to deliver sustainable design patterns
- Drive early adoption of non-functional production support requirements
- Automate services to improve reliability and efficiency
- Serve as a senior technical authority for GCP platform reliability, resiliency, observability, automation, and production readiness
- Design and implement reliability patterns for Azure landing zones, private networking, DNS, firewalls, regional resiliency, service health checks, and workload onboarding
- Lead platform reliability initiatives including secondary-region readiness, ingress/egress observability, DNS resolver monitoring, GenAI platform health checks, and enterprise dashboard automation
- Define and mature SLIs, SLOs, reliability indicators, alerting standards, and service health reporting
- Develop Terraform modules, automation frameworks, and CI/CD patterns
- Drive observability improvements using Log Analytics, Dynatrace, Resource Graph, dashboards, and enterprise monitoring tools
- Identify reliability risks and translate them into roadmaps, remediation plans, automation opportunities, and operational controls
- Lead technical investigations for major incidents, recurring problems, platform defects, and service degradation
- Partner with security and governance teams on IAM, policy-as-code, vulnerability remediation, control validation, and audit readiness
- Provide technical design input for new Azure services and workloads before production adoption
- Mentor SRE engineers and establish standards for runbooks, dashboards, health checks, problem records, post-incident reviews, and production-readiness gates
- Visualize production support metrics for operational readiness and SRE teams
- Develop software solutions and processes to eliminate toil
- Create error budget policies and recommend code optimizations, instrumentation, and logging to improve reliability visibility
- Plan for capacity bottlenecks, vulnerabilities, error rates, and reliability improvements
- Assess monitoring for new changes and enhance application and system monitoring designs
- Serve as a subject matter expert in incident triage, failure scenario modeling, and complex problem management investigations
- Collaborate with Development and Infrastructure teams to develop SLIs and SLOs
Requirements
What you’ll need- 4+ years of experience in cloud infrastructure engineering, platform engineering, or cloud operations, with exposure to Google Cloud Platform (GCP)
- Strong hands-on experience with Infrastructure as Code (IaC), including Terraform or Terraform Enterprise
- Understanding of software engineering fundamentals, version control, code quality, and basic testing practices for infrastructure code
- Experience developing and maintaining Terraform modules and infrastructure configurations
- Familiarity with CI/CD pipelines for infrastructure deployment
- Working knowledge of DevSecOps practices, including security and compliance checks in automated workflows
- Understanding of GCP services and cloud architecture fundamentals, including VPCs, IAM, and load balancing
- Exposure to policy-as-code, governance, and compliance requirements in enterprise environments
- Experience supporting automation and standardization efforts in cloud deployments
- Understanding of monitoring, logging, and observability tools
- Hands-on experience with incident response, troubleshooting, and root cause analysis in cloud or distributed systems
- Ability to collaborate with engineering, architecture, and security teams
- Strong problem-solving and analytical skills
- Effective communication skills and ability to document technical solutions clearly
- Interest in emerging technologies and automation techniques, including AI/ML where applicable
- Availability for 1st shift in the United States of America, 40 hours per week
Benefits
Comp & perks- Discretionary incentive eligible; eligible to participate in the annual discretionary plan
- Annual discretionary award based on individual performance, line of business/group performance, and overall Company success
- Benefits eligible
- Access to paid time off
- Resources and support to employees
- Opportunities to learn, grow, and make an impact
- Inclusive workplace
- Support for physical, emotional, and financial wellness
- Recognition and rewards for performance