Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
KeyBank

Resiliency Engineer

KeyBank

. Design, build, and continuously improve the reliability, availability, and recoverability of KeyBank technology platforms across on-premises, hybrid, and cloud environments .

Posted 9/24/2026full-timeAlbany • New York • United StatesMid-LevelSenior💰 $63,000 - $96,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in reliability engineering, automation, and infrastructure-as-code, with a strong focus on disaster recovery planning and service-level management. Proficient in coding and implementing solutions across on-premises and cloud environments, ensuring high availability and compliance with regulatory standards.

Highest-signal resume keywords
Reliability EngineeringInfrastructure-As-CodeDisaster Recovery PlanningAutomation Using PythonService-Level Management

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Reliability EngineeringInfrastructure-As-CodeDisaster Recovery PlanningAutomation Using PythonHigh-Availability DesignSite Reliability Engineering PrinciplesChaos/Fault-Injection EngineeringMonitoring and ObservabilityHealth ChecksSLIs and SLOs
Soft Skills
Analytical SkillsProblem-Solving SkillsWritten CommunicationVerbal CommunicationFacilitation Skills
Tools & Technologies
TerraformAnsibleGCPAzureServiceNowCI/CDContainersOrchestration PlatformsMonitoring ToolsObservability Tooling
Certifications & Qualifications
Cloud ArchitectCloud EngineerKubernetesLinuxITIL
Industry Keywords
Regulated IndustryFFIECNIST SP 800-34NIST CSFISO 22301

Tech Stack

Tools & technologies
AnsibleAzureCloudGoogle Cloud PlatformKubernetesLinuxPythonServiceNowTerraform

About the role

Key responsibilities & impact
  • Design, build, and continuously improve the reliability, availability, and recoverability of KeyBank technology platforms across on-premises, hybrid, and cloud environments
  • Design and code automation to reduce operational toil and replace manual runbooks with orchestrated, auditable failover and recovery workflows
  • Develop and maintain infrastructure-as-code, scripts, and pipelines to provision, configure, and validate recovery environments
  • Build self-healing patterns, health checks, and automated validation to confirm recoverability
  • Partner with application and infrastructure teams to assess architecture for reliability, redundancy, and recoverability
  • Define, measure, and support SLIs, SLOs, and error budgets for critical services
  • Ensure architecture meets RTO and RPO targets
  • Plan and execute disaster recovery tests and targeted fault-injection/chaos experiments
  • Improve monitoring, alerting, and observability to detect service degradation early
  • Facilitate resiliency and architecture reviews, tabletop exercises, and cross-team recovery walkthroughs
  • Provide subject-matter expertise on reliability engineering practices and drive adoption across technology teams
  • Produce examiner-ready documentation and evidence aligned with KeyBank policies, standards, and regulatory requirements

Requirements

What you’ll need
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or related field—or equivalent work experience
  • Demonstrated experience in reliability engineering, DevOps, infrastructure, or technology operations
  • Hands-on coding and automation ability using tools such as Python, Bash, or PowerShell
  • Experience with infrastructure-as-code such as Terraform or Ansible
  • Working knowledge of both on-premises infrastructure and public cloud platforms, including GCP and/or Azure
  • Understanding of high-availability design, including redundancy, replication, failover, and load balancing
  • Strong facilitation, analytical, problem-solving, and written/verbal communication skills
  • Experience with site reliability engineering principles and service-level management
  • Experience with disaster recovery planning, resiliency testing, or chaos/fault-injection engineering
  • Familiarity with containers and orchestration, CI/CD, and observability tooling
  • Experience with ServiceNow or comparable orchestration platforms
  • Experience in a regulated industry or large, complex enterprise environment
  • Familiarity with FFIEC, NIST SP 800-34/CSF, or ISO 22301
  • Relevant certifications such as cloud architect/engineer, Kubernetes, Linux, or ITIL

Benefits

Comp & perks
  • Base salary range of $63,000.00–$96,000.00 annually
  • Eligibility for incentive compensation, which may include production, commission, and/or discretionary incentives
  • Benefits for which the position is eligible
  • Flexible options in circumstances where roles can be performed effectively in a mobile environment