FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Resiliency Engineer
KeyBank. Design, build, and continuously improve the reliability, availability, and recoverability of KeyBank technology platforms across on-premises, hybrid, and cloud environments .
Posted 9/24/2026full-timeAlbany • New York • United StatesMid-LevelSenior💰 $63,000 - $96,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in reliability engineering, automation, and infrastructure-as-code, with a strong focus on disaster recovery planning and service-level management. Proficient in coding and implementing solutions across on-premises and cloud environments, ensuring high availability and compliance with regulatory standards.
Highest-signal resume keywords
Reliability EngineeringInfrastructure-As-CodeDisaster Recovery PlanningAutomation Using PythonService-Level Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Reliability EngineeringInfrastructure-As-CodeDisaster Recovery PlanningAutomation Using PythonHigh-Availability DesignSite Reliability Engineering PrinciplesChaos/Fault-Injection EngineeringMonitoring and ObservabilityHealth ChecksSLIs and SLOs
Soft Skills
Analytical SkillsProblem-Solving SkillsWritten CommunicationVerbal CommunicationFacilitation Skills
Tools & Technologies
TerraformAnsibleGCPAzureServiceNowCI/CDContainersOrchestration PlatformsMonitoring ToolsObservability Tooling
Certifications & Qualifications
Cloud ArchitectCloud EngineerKubernetesLinuxITIL
Industry Keywords
Regulated IndustryFFIECNIST SP 800-34NIST CSFISO 22301
Tech Stack
Tools & technologiesAnsibleAzureCloudGoogle Cloud PlatformKubernetesLinuxPythonServiceNowTerraform
About the role
Key responsibilities & impact- Design, build, and continuously improve the reliability, availability, and recoverability of KeyBank technology platforms across on-premises, hybrid, and cloud environments
- Design and code automation to reduce operational toil and replace manual runbooks with orchestrated, auditable failover and recovery workflows
- Develop and maintain infrastructure-as-code, scripts, and pipelines to provision, configure, and validate recovery environments
- Build self-healing patterns, health checks, and automated validation to confirm recoverability
- Partner with application and infrastructure teams to assess architecture for reliability, redundancy, and recoverability
- Define, measure, and support SLIs, SLOs, and error budgets for critical services
- Ensure architecture meets RTO and RPO targets
- Plan and execute disaster recovery tests and targeted fault-injection/chaos experiments
- Improve monitoring, alerting, and observability to detect service degradation early
- Facilitate resiliency and architecture reviews, tabletop exercises, and cross-team recovery walkthroughs
- Provide subject-matter expertise on reliability engineering practices and drive adoption across technology teams
- Produce examiner-ready documentation and evidence aligned with KeyBank policies, standards, and regulatory requirements
Requirements
What you’ll need- Bachelor’s degree in Computer Science, Information Technology, Engineering, or related field—or equivalent work experience
- Demonstrated experience in reliability engineering, DevOps, infrastructure, or technology operations
- Hands-on coding and automation ability using tools such as Python, Bash, or PowerShell
- Experience with infrastructure-as-code such as Terraform or Ansible
- Working knowledge of both on-premises infrastructure and public cloud platforms, including GCP and/or Azure
- Understanding of high-availability design, including redundancy, replication, failover, and load balancing
- Strong facilitation, analytical, problem-solving, and written/verbal communication skills
- Experience with site reliability engineering principles and service-level management
- Experience with disaster recovery planning, resiliency testing, or chaos/fault-injection engineering
- Familiarity with containers and orchestration, CI/CD, and observability tooling
- Experience with ServiceNow or comparable orchestration platforms
- Experience in a regulated industry or large, complex enterprise environment
- Familiarity with FFIEC, NIST SP 800-34/CSF, or ISO 22301
- Relevant certifications such as cloud architect/engineer, Kubernetes, Linux, or ITIL
Benefits
Comp & perks- Base salary range of $63,000.00–$96,000.00 annually
- Eligibility for incentive compensation, which may include production, commission, and/or discretionary incentives
- Benefits for which the position is eligible
- Flexible options in circumstances where roles can be performed effectively in a mobile environment