Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Bank of America

Senior Site Reliability Engineer

Bank of America

. Partner with engineering and technology leaders to define objective reliability goals for services .

Posted 9/18/2026full-timeJersey City • New Jersey • United StatesSenior💰 $152,600 - $191,500 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in cloud infrastructure engineering and platform reliability, with a strong focus on Google Cloud Platform (GCP) and Terraform for automation and infrastructure as code. Capable of leading reliability initiatives, defining service level indicators, and enhancing observability and operational readiness.

Highest-signal resume keywords
Google Cloud Platform (GCP)Infrastructure as Code (IaC)Terraform DevelopmentCI/CD Pipeline ManagementObservability Tools

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Cloud Infrastructure EngineeringPlatform EngineeringAutomation FrameworksService Level Indicators (SLIs)Service Level Objectives (SLOs)Incident ResponseRoot Cause AnalysisPolicy-as-CodeDevSecOps PracticesMonitoring and Logging
Soft Skills
Effective CommunicationProblem-SolvingAnalytical SkillsCollaboration
Tools & Technologies
TerraformLog AnalyticsDynatraceResource GraphCI/CD Tools
Industry Keywords
Cloud OperationsResiliency DesignAutomationGovernanceCompliance

Tech Stack

Tools & technologies
AzureCloudDistributed SystemsDNSFirewallsGoogle Cloud PlatformTerraform

About the role

Key responsibilities & impact
  • Partner with engineering and technology leaders to define objective reliability goals for services
  • Serve as a senior technical authority for GCP platform reliability, resiliency, observability, automation, and production readiness
  • Design and implement advanced reliability patterns for Azure landing zones, private networking, DNS, firewalls, regional resiliency, service health checks, and workload onboarding
  • Lead complex platform reliability initiatives, including secondary-region readiness, egress/ingress observability, private DNS resolver monitoring, GenAI platform health checks, and enterprise dashboard automation
  • Define and mature SLIs, SLOs, reliability indicators, alerting standards, and service health reporting
  • Develop reusable Terraform modules, automation frameworks, and CI/CD patterns
  • Drive observability improvements using Log Analytics, Dynatrace, Resource Graph, dashboards, and enterprise monitoring tools
  • Identify systemic reliability risks and translate them into engineering roadmaps, remediation plans, automation opportunities, and operational controls
  • Lead technical investigations for major incidents, recurring problems, platform defects, and service degradation events
  • Integrate IAM, policy-as-code, vulnerability remediation, control validation, and audit readiness into Azure platform operations
  • Provide technical design input for new Azure services and workloads to ensure operational readiness
  • Mentor SRE engineers and raise standards for automation, troubleshooting, documentation, resiliency design, and production support
  • Create executive-ready technical summaries, reliability narratives, and leadership recommendations
  • Establish standards for runbooks, dashboards, health checks, problem records, post-incident reviews, and production-readiness gates
  • Visualize key production support metrics to identify scenarios requiring intervention
  • Develop software solutions and improved processes to reduce toil and manual support effort
  • Create error budget policies and improve service reliability visibility through code optimization, instrumentation, and logging
  • Plan for capacity bottlenecks, vulnerabilities, error rates, and other reliability improvement opportunities
  • Assess monitoring for new changes and enhance application and system monitoring designs
  • Engage as a subject matter expert in incident triage, failure scenario modeling, and root-cause investigations
  • Collaborate with Development and Infrastructure teams to develop SLIs and SLOs

Requirements

What you’ll need
  • 4+ years of experience in cloud infrastructure engineering, platform engineering, or cloud operations, with exposure to Google Cloud Platform (GCP)
  • Strong hands-on experience with Infrastructure as Code (IaC), including practical use of Terraform or Terraform Enterprise
  • Solid understanding of software engineering fundamentals, including version control, code quality, and basic testing practices for infrastructure code
  • Experience developing and maintaining Terraform modules and infrastructure configurations
  • Familiarity with CI/CD pipelines for infrastructure deployment, including automated build, test, and release processes
  • Working knowledge of DevSecOps practices, including integrating security and compliance checks into automated workflows
  • Good understanding of GCP services and cloud architecture fundamentals, including VPCs, IAM, and load balancing
  • Exposure to policy-as-code, governance, and compliance requirements in enterprise environments
  • Experience supporting automation and standardization efforts in cloud deployments
  • Understanding of monitoring, logging, and observability tools
  • Hands-on experience with incident response, troubleshooting, and root cause analysis in cloud or distributed systems
  • Ability to collaborate with engineering, architecture, and security teams
  • Strong problem-solving and analytical skills
  • Effective communication skills and ability to document technical solutions clearly
  • Interest in emerging technologies and automation techniques, including AI/ML where applicable

Benefits

Comp & perks
  • Discretionary incentive eligible
  • Eligible to participate in the annual discretionary plan
  • Annual discretionary award based on individual performance, line of business/group performance, and overall Company success
  • Benefits eligible
  • Access to paid time off
  • Resources and support to employees
  • Opportunities to learn, grow, and make an impact
  • Inclusive workplace
  • Support for physical, emotional, and financial wellness
  • Recognition and reward for performance