FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Site Reliability Engineering (SRE) practices, including automation, performance tuning, and incident management. Capable of leading teams, mentoring engineers, and ensuring high availability and resilience of production systems.
Highest-signal resume keywords
Site Reliability Engineering (SRE)Performance TuningIncident ManagementAutomationTeam Leadership
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Error BudgetsSLIs/SLOsCapacity PlanningJVM TuningThread Pool OptimizationConnection Pool ManagementMiddleware TroubleshootingLog-Based MonitoringAnomaly DetectionITIL Compliance
Soft Skills
Analytical SkillsTroubleshooting SkillsClear CommunicationCollaborationCalm Under Pressure
Tools & Technologies
WebLogicTomcatApacheNginxNew RelicSplunkAkamai
Industry Keywords
High AvailabilityPerformanceResilienceChaos TestingRunbook CreationPost-Incident ReviewsCross-Functional Collaboration
Tech Stack
Tools & technologiesApacheNGINXSplunk
About the role
Key responsibilities & impact- Ensure high availability, performance, and resilience of production systems.
- Implement SRE best practices, including error budgets, SLIs/SLOs, capacity planning, chaos testing, and runbook creation.
- Drive automation to reduce manual operational tasks and improve MTTR.
- Conduct post-incident reviews and implement long-term corrective actions.
- Manage deployments, rollbacks, and environment synchronization across Dev, QA, UAT, and Production.
- Install, configure, upgrade, and maintain WebLogic, Tomcat, Apache, and Nginx servers.
- Perform JVM tuning, thread pool optimization, connection pool management, and performance tuning.
- Troubleshoot middleware issues involving memory leaks, thread contention, SSL, certificates, and clustering.
- Configure dashboards, alerts, and performance insights using New Relic and Splunk.
- Develop log-based monitoring strategies and anomaly detection.
- Configure Akamai caching rules, WAF policies, edge redirects, and performance optimizations.
- Troubleshoot CDN latency, caching, and routing issues; collaborate with Akamai support.
- Lead major incident bridges and coordinate cross-functional teams.
- Manage problem tickets, root cause analysis, and preventive action plans.
- Ensure compliance with ITIL processes for change, release, and incident management.
- Lead and mentor a team of SRE/DevOps engineers.
- Provide technical guidance, training, and performance feedback.
- Act as a customer-facing technical SME for escalations and production issues.
- Collaborate with product, QA, development, and business teams to ensure smooth delivery.
Requirements
What you’ll need- Experience, education, certifications, and technical prerequisites are not specified in the posting.
- Ability to lead and mentor SRE/DevOps engineers.
- Strong analytical and troubleshooting skills for complex issues.
- Clear, structured communication with customers and internal teams.
- Ability to collaborate across engineering, QA, product, and business teams.
- Ability to handle critical incidents calmly and clearly.
