FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates extensive experience in Site Reliability Engineering for SaaS applications, focusing on improving reliability, performance, and scalability in cloud environments. Proficient in automation, monitoring, and incident management to ensure operational excellence.
Highest-signal resume keywords
Site Reliability EngineeringMicrosoft AzureAutomation in PowerShell, Python, BashSLIs, SLOs, and Error BudgetsIncident Response and Root Cause Analysis
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Site Reliability EngineeringAutomation in PowerShellAutomation in PythonAutomation in BashInfrastructure as CodeMonitoring and LoggingPerformance TestingCapacity PlanningTroubleshootingDisaster Recovery
Soft Skills
CommunicationCollaborationTechnical Documentation
Tools & Technologies
DatadogAzure Application InsightsElasticTerraformAnsibleAzure DevOpsCloudflare
Industry Keywords
SaaSPublic CloudOperational SecurityCompliance InitiativesCI/CD Pipelines
Tech Stack
Tools & technologiesAnsibleAzureCloudFirewallsPythonSQLTerraform.NET
About the role
Key responsibilities & impact- Perform Site Reliability Engineering for JustFOIA, MCCi’s SaaS platform for government public-records requests
- Build and maintain monitoring, alerting, dashboards, logging, and operational reporting across application, database, and infrastructure tiers
- Analyze logs, metrics, traces, and service-health indicators to identify emerging and recurring issues
- Improve platform reliability, availability, performance, scalability, and maintainability
- Participate in on-call rotation and lead incident triage, troubleshooting, escalation, and service restoration
- Isolate production issues across application, IIS/.NET, SQL Server, network, and cloud-platform components
- Lead blameless post-incident reviews, complete root cause analyses, and drive corrective actions
- Operate and improve Azure and Azure Government environments
- Develop automation to reduce repetitive work, human error, and configuration inconsistency
- Maintain and enhance Infrastructure-as-Code, configuration management, deployment, and operational tooling
- Support CI/CD pipelines and application release processes
- Identify performance bottlenecks, capacity risks, and scalability limitations
- Define and track service-level indicators, objectives, error budgets, or equivalent reliability measures
- Support backup validation, disaster recovery, failover testing, capacity planning, and operational readiness reviews
- Support vulnerability remediation, operational security controls, and compliance initiatives
- Create and maintain runbooks, troubleshooting guides, recovery procedures, and operational documentation
- Partner with development teams to improve application reliability, observability, and supportability
- Contribute to deployment automation, infrastructure improvements, and evolving DevOps practices
- Potentially support other MCCi SaaS solutions as Cloud Operations responsibilities evolve
Requirements
What you’ll need- Six or more years of overall IT experience
- Three or more years performing site reliability engineering for production SaaS workloads in a public cloud, with direct responsibility for availability and performance
- Hands-on production Site Reliability Engineering experience for a SaaS application
- Experience defining and operating against SLIs, SLOs, and error budgets
- Experience building observability with metrics, centralized logging, distributed tracing, dashboards, and application performance monitoring
- Experience designing actionable alerting and reducing alert noise
- Experience leading incident response, root cause analysis, and blameless post-incident reviews
- Experience driving reliability improvements to completion
- Experience writing automation in PowerShell, Python, Bash, or a comparable language
- Experience with Infrastructure as Code and configuration management
- Experience operating workloads in Microsoft Azure or another major public cloud
- Experience supporting customer-facing, business-critical applications
- Strong troubleshooting, communication, collaboration, and technical documentation skills
- Preferred: Windows Server and IIS/.NET application support in production
- Preferred: SQL Server monitoring and performance troubleshooting
- Preferred: CI/CD pipelines and deployment automation
- Preferred: Performance testing and capacity planning
- Preferred: Network traffic management, load balancing, web application firewalls, and content delivery
- Preferred: Backup, replication, failover, disaster recovery, and recovery testing
- Preferred: Containers and orchestration technologies
- Experience with Datadog, Azure Application Insights, Elastic, Terraform, Ansible, PowerShell, Azure DevOps, or Cloudflare is a plus
Benefits
Comp & perks- Remote work and virtual inclusion for remote teammates
- Approachable management and leadership
- Trust-based work culture with minimal rules
- Comfortable dress code
- Inclusive, diverse, and respectful workplace
- Team collaboration, relationship building, and recognition of good work
