Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
BMO U.S.

Senior SRE Engineer – CCaaS

BMO U.S.

. Ensure 24x7 availability, reliability, performance, security and operational excellence of the CCaaS platform and supporting services .

Posted 9/15/2026full-timeRemote • CanadaSenior💰 CA$70,000 - CA$150,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in Site Reliability Engineering (SRE) practices, including incident management, automation, and performance optimization. Proficient in utilizing observability tools and cloud services to ensure platform reliability and operational excellence.

Highest-signal resume keywords
Site Reliability Engineering (SRE)Automation and Automation PipelinesIncident ManagementCloud ComputingMonitoring and Alerting Configuration

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
PythonPowerShellAWS LambdaCloudWatchDynatraceOpenSearchRobot Process AutomationConfiguration ManagementContainer OrchestrationSystem Design
Soft Skills
Verbal CommunicationWritten CommunicationCollaborationAnalytical SkillsProblem-Solving Skills
Tools & Technologies
SplunkGrafanaAPM ToolsMonitoring DashboardsHealth Checks
Industry Keywords
CCaaSOperational ExcellenceDisaster RecoveryCapacity PlanningSecurity Compliance

Tech Stack

Tools & technologies
AWSCloudCyber SecurityGrafanaPythonSplunk

About the role

Key responsibilities & impact
  • Ensure 24x7 availability, reliability, performance, security and operational excellence of the CCaaS platform and supporting services
  • Define and manage SLIs, SLOs and error budgets
  • Monitor platform health, performance, capacity and stability
  • Build and maintain monitoring dashboards, alerts and health checks
  • Configure and optimize CloudWatch, Splunk, Dynatrace, Grafana and other observability tools
  • Develop operational runbooks and troubleshooting guides
  • Act as an escalation point for critical incidents and service disruptions
  • Lead incident triage, impact assessment, communication and recovery activities
  • Perform root cause analysis and identify preventive actions
  • Drive reduction of MTTR, MTTD and recurring incidents
  • Eliminate operational toil through automation and develop self-healing and automated remediation solutions
  • Build scripts and tools using Python, PowerShell, AWS Lambda and other automation frameworks
  • Automate operational checks, deployment validation, monitoring and reporting
  • Support Amazon Connect, AWS services, APIs, middleware, routing, IVR and integration components
  • Manage platform configurations, certificates, service accounts and connectivity requirements
  • Support disaster recovery, backup, restoration and failover testing
  • Conduct capacity planning and performance optimization
  • Participate in release planning, deployment validation and production implementation
  • Review changes for operational risk and reliability impact
  • Support production deployments and rollback planning
  • Ensure operational readiness before launch
  • Support vulnerability remediation, audits, security events and compliance requirements
  • Analyze operational metrics and identify improvement opportunities
  • Drive SRE best practices across CCaaS teams
  • Participate in architecture reviews focused on scalability and resiliency
  • Mentor developers and operations teams on SRE practices
  • Partner with Development, DevOps, Infrastructure, Security, Business and Vendor teams
  • Debug production issues across services and technology-stack levels
  • Improve service health visibility through metrics, logs and traces
  • Compute the cost and business impact of SLA breaches and downtime

Requirements

What you’ll need
  • Typically 5–6 years of relevant experience
  • Post-secondary degree in a related field of study or an equivalent combination of education and experience
  • Deep knowledge and technical proficiency gained through extensive education and business experience
  • Foundational proficiency in DevOps
  • Foundational proficiency in cybersecurity and privacy concepts, principles and solutions
  • Intermediate proficiency in IT infrastructure library
  • Intermediate proficiency in Robot Process Automation
  • Intermediate proficiency in cloud computing
  • Intermediate proficiency in configuration management
  • Intermediate proficiency in container orchestration
  • Intermediate proficiency in system design and implementation
  • Intermediate proficiency in incident management
  • Advanced proficiency in alerting and log configuration with OpenSearch and CloudWatch
  • Advanced proficiency in Dynatrace or another APM tool configuration and dashboarding
  • Advanced proficiency in automation and automation pipelines
  • Advanced proficiency in automated testing
  • Proficiency in verbal and written communication, collaboration, analytical and problem-solving skills, and data-driven decision making

Benefits

Comp & perks
  • Performance-based incentives
  • Discretionary bonuses
  • Health insurance
  • Tuition reimbursement
  • Accident and life insurance
  • Retirement savings plans
  • In-depth training and coaching
  • Manager support
  • Network-building opportunities
  • Accommodations available on request during the selection process