FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing, deploying, and operating enterprise observability platforms, with a strong focus on Splunk, Elasticsearch, and distributed tracing technologies. Proficient in automation and infrastructure as code practices, ensuring high availability and reliability in production environments.
Highest-signal resume keywords
Splunk Enterprise AdministrationElasticsearch Platform DesignInfrastructure As Code (Terraform)Distributed Tracing (Grafana Tempo, OpenTelemetry)Production Incident Troubleshooting
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Splunk SPLElasticsearchPrometheusGrafanaKafkaPythonGoRubyBashTerraform
Soft Skills
CommunicationCollaborationProblem-Solving
Tools & Technologies
Splunk CloudGrafana TempoKubernetesAWSAzureGCPAnsibleCI/CD PipelinesService Mesh TechnologiesConsul
Certifications & Qualifications
Splunk Certification
Industry Keywords
Site Reliability EngineeringPlatform EngineeringDevOpsObservability StrategiesCapacity PlanningIncident ManagementFedRAMPRegulated EnvironmentsHigh AvailabilityOperational Support
Tech Stack
Tools & technologiesAnsibleAWSAzureCloudConsulDistributed SystemsElasticSearchGoogle Cloud PlatformGrafanaKafkaKubernetesLinuxPrometheusPythonRubySplunkTerraformGo
About the role
Key responsibilities & impact- Design, deploy, operate, and continuously improve enterprise observability platforms
- Build and maintain Splunk Enterprise and Splunk Cloud infrastructure, including indexers, Search Head Clusters, heavy forwarders, deployment servers, and related integrations
- Deploy and operate large-scale Elasticsearch clusters for log analytics, search, and operational troubleshooting
- Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry
- Build and maintain end-to-end telemetry pipelines for logs, metrics, and traces
- Define instrumentation standards, data quality practices, retention policies, and observability patterns
- Scale and optimize Prometheus, Grafana, Kafka, Tempo, OpenTelemetry, and related monitoring systems
- Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo
- Establish monitoring and alerting standards to improve service-level visibility and reduce mean time to detect and resolve incidents
- Automate infrastructure provisioning, configuration, upgrades, and operational processes using Terraform and configuration management tools
- Participate in capacity planning, performance tuning, upgrades, patching, disaster recovery, and production readiness reviews
- Troubleshoot complex distributed-systems issues and lead or support incident response and root-cause analysis
- Partner with platform, application, security, database, and network engineering teams to improve reliability and operational consistency
- Contribute to shared SRE standards, documentation, runbooks, and engineering best practices
- Participate in the shared SRE pager and on-call rotation for production services
- Support customer-facing production services and Kubernetes-based platform workloads such as Nextunnel
- Improve visibility, reduce time to detect and resolve incidents, increase platform reliability, and strengthen operational excellence across Cisco’s cloud environment
Requirements
What you’ll need- 7+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or a related field
- Experience administering Splunk Enterprise or Splunk Cloud in a production environment
- Solid understanding of Splunk SPL and experience developing production dashboards, alerts, and analytics
- Experience designing, deploying, and operating Elasticsearch or ELK-based platforms
- Experience with Prometheus, Grafana, Grafana Tempo, OpenTelemetry, distributed tracing, and Kafka
- Demonstrated experience implementing and operating modern observability strategies across logs, metrics, and traces
- Experience with Terraform and Infrastructure as Code practices
- Strong understanding of Linux, networking, distributed systems, high availability, and production operations
- Programming or scripting experience in Python, Go, Ruby, Bash, or a similar language
- Experience troubleshooting production incidents and working within an on-call or operational support model
- Strong communication, collaboration, and problem-solving skills
- Must be a U.S. Person (U.S. citizen or U.S. national)
- Work must be performed from within the United States
- Preferred: Splunk certification
- Preferred: Experience with Kubernetes and containerized production workloads
- Preferred: Experience operating cloud infrastructure in AWS, Azure, or GCP
- Preferred: Experience with Ansible, Consul, CI/CD pipelines, or service mesh technologies
- Preferred: Experience with SLOs, SLIs, error budgets, capacity planning, and incident management
- Preferred: Experience supporting FedRAMP, government, or other regulated environments
- Preferred: Experience building secure, highly available, and compliant observability platforms
- Preferred: Experience leading cross-functional technical initiatives from proposal through implementation
Benefits
Comp & perks- Medical, dental and vision insurance
- 401(k) plan with a Cisco matching contribution
- Paid parental leave
- Short- and long-term disability coverage
- Basic life insurance
- Cisco restricted stock unit grants may be available
- 10 paid holidays per full calendar year
- 1 floating holiday for non-exempt employees
- 1 paid day off for employee’s birthday
- Paid year-end holiday shutdown
- 4 paid days off for personal wellness
- 16 days of paid vacation per full calendar year for non-exempt employees
- Flexible vacation time off program with no defined limit for eligible exempt employees
- 80 hours of sick time off provided on hire date and each January 1st thereafter
- Up to 80 hours of unused sick time carried forward annually
- Additional paid time away for critical or emergency family issues
- Optional 10 paid volunteer days per full calendar year
- Annual bonuses for non-sales roles, subject to Cisco policies
