Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Bank of America

Senior Site Reliability Engineer

Bank of America

. Partner with engineering and technology leaders to define objective reliability goals for services .

Posted 9/18/2026full-timeJersey City • New Jersey • United StatesSenior💰 $152,600 - $191,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in cloud infrastructure engineering and platform reliability, with a strong focus on Google Cloud Platform (GCP) and automation practices. Proficient in designing observability solutions, implementing Infrastructure as Code (IaC) with Terraform, and leading cross-functional initiatives to enhance service reliability and operational readiness.

Highest-signal resume keywords
Google Cloud Platform (GCP)Infrastructure as Code (IaC)Terraform DevelopmentObservability Solutions DesignIncident Response and Root Cause Analysis

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Cloud Infrastructure EngineeringPlatform EngineeringTerraform Modules DevelopmentCI/CD PipelinesMonitoring and Logging ToolsDevSecOps PracticesService Level Indicators (SLIs)Service Level Objectives (SLOs)Automation FrameworksError Budget Policies
Soft Skills
Problem-Solving SkillsEffective CommunicationCollaboration with Cross-Functional TeamsAnalytical SkillsMentoring and Leadership
Tools & Technologies
Log AnalyticsDynatraceResource GraphEnterprise Monitoring ToolsPolicy-as-Code
Industry Keywords
Cloud OperationsGovernance and ComplianceIncident TriageService Health ChecksAutomation Opportunities

Tech Stack

Tools & technologies
AzureCloudDistributed SystemsDNSFirewallsGoogle Cloud PlatformTerraform

About the role

Key responsibilities & impact
  • Partner with engineering and technology leaders to define objective reliability goals for services
  • Design observability solutions through instrumentation and dashboards
  • Identify root causes of complex and high-impact issues
  • Partner with cross-functional teams to deliver sustainable design patterns
  • Drive early adoption of non-functional production support requirements
  • Automate services to improve reliability and efficiency
  • Serve as a senior technical authority for GCP platform reliability, resiliency, observability, automation, and production readiness
  • Design and implement reliability patterns for Azure landing zones, private networking, DNS, firewalls, regional resiliency, service health checks, and workload onboarding
  • Lead platform reliability initiatives including secondary-region readiness, ingress/egress observability, DNS resolver monitoring, GenAI platform health checks, and enterprise dashboard automation
  • Define and mature SLIs, SLOs, reliability indicators, alerting standards, and service health reporting
  • Develop Terraform modules, automation frameworks, and CI/CD patterns
  • Drive observability improvements using Log Analytics, Dynatrace, Resource Graph, dashboards, and enterprise monitoring tools
  • Identify reliability risks and translate them into roadmaps, remediation plans, automation opportunities, and operational controls
  • Lead technical investigations for major incidents, recurring problems, platform defects, and service degradation
  • Partner with security and governance teams on IAM, policy-as-code, vulnerability remediation, control validation, and audit readiness
  • Provide technical design input for new Azure services and workloads before production adoption
  • Mentor SRE engineers and establish standards for runbooks, dashboards, health checks, problem records, post-incident reviews, and production-readiness gates
  • Visualize production support metrics for operational readiness and SRE teams
  • Develop software solutions and processes to eliminate toil
  • Create error budget policies and recommend code optimizations, instrumentation, and logging to improve reliability visibility
  • Plan for capacity bottlenecks, vulnerabilities, error rates, and reliability improvements
  • Assess monitoring for new changes and enhance application and system monitoring designs
  • Serve as a subject matter expert in incident triage, failure scenario modeling, and complex problem management investigations
  • Collaborate with Development and Infrastructure teams to develop SLIs and SLOs

Requirements

What you’ll need
  • 4+ years of experience in cloud infrastructure engineering, platform engineering, or cloud operations, with exposure to Google Cloud Platform (GCP)
  • Strong hands-on experience with Infrastructure as Code (IaC), including Terraform or Terraform Enterprise
  • Understanding of software engineering fundamentals, version control, code quality, and basic testing practices for infrastructure code
  • Experience developing and maintaining Terraform modules and infrastructure configurations
  • Familiarity with CI/CD pipelines for infrastructure deployment
  • Working knowledge of DevSecOps practices, including security and compliance checks in automated workflows
  • Understanding of GCP services and cloud architecture fundamentals, including VPCs, IAM, and load balancing
  • Exposure to policy-as-code, governance, and compliance requirements in enterprise environments
  • Experience supporting automation and standardization efforts in cloud deployments
  • Understanding of monitoring, logging, and observability tools
  • Hands-on experience with incident response, troubleshooting, and root cause analysis in cloud or distributed systems
  • Ability to collaborate with engineering, architecture, and security teams
  • Strong problem-solving and analytical skills
  • Effective communication skills and ability to document technical solutions clearly
  • Interest in emerging technologies and automation techniques, including AI/ML where applicable
  • Availability for 1st shift in the United States of America, 40 hours per week

Benefits

Comp & perks
  • Discretionary incentive eligible; eligible to participate in the annual discretionary plan
  • Annual discretionary award based on individual performance, line of business/group performance, and overall Company success
  • Benefits eligible
  • Access to paid time off
  • Resources and support to employees
  • Opportunities to learn, grow, and make an impact
  • Inclusive workplace
  • Support for physical, emotional, and financial wellness
  • Recognition and rewards for performance