FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and maintaining enterprise monitoring and observability solutions, with a strong focus on Grafana, IBM Instana, and telemetry data analysis. Proficient in automation using Python, PowerShell, and Infrastructure as Code methodologies.
Highest-signal resume keywords
Grafana DevelopmentIBM Instana AdministrationPrometheus MonitoringAnsible AutomationIncident Management
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
GrafanaIBM InstanaSolarWindsTelegrafPrometheusInfluxDBPythonPowerShellLinux Shell ScriptingVBScript
Soft Skills
TroubleshootingRoot Cause AnalysisOperational SupportCollaborationCommunication
Tools & Technologies
VMwareCitrix VDIServiceNowITIL FrameworkELT Platforms
Industry Keywords
ObservabilitySLO MonitoringSLA MonitoringInfrastructure OperationsSite Reliability Engineering
Tech Stack
Tools & technologiesAnsibleAWSAzureCitrixCloudDNSDockerGoogle Cloud PlatformGrafanaIoTJenkinsKubernetesLinuxMicroservicesPrometheusPuppetPythonServiceNowShell ScriptingVMware
About the role
Key responsibilities & impact- Design, implement, and maintain enterprise monitoring and observability solutions
- Develop and maintain Grafana dashboards, alerts, and visualizations
- Monitor infrastructure, applications, middleware, IoT services, and enterprise telemetry using IBM Instana, Grafana, SolarWinds, and related tools
- Configure and manage data collection with Telegraf, Prometheus, and monitoring agents
- Analyze metrics, logs, traces, events, and telemetry to identify bottlenecks and service degradation
- Support SLO, SLA, and operational health monitoring initiatives
- Perform root cause analysis and troubleshoot infrastructure and application issues
- Support onboarding, monitoring, and operational management of FOAK services and enterprise applications
- Configure, validate, and troubleshoot ELT integrations and telemetry pipelines
- Monitor log ingestion, event correlation, and telemetry data quality
- Monitor Linux and Windows servers, VMware, Citrix VDI, DNS, proxy, middleware, integration services, enterprise applications, and IoT platforms
- Investigate performance issues, recurring alerts, infrastructure anomalies, capacity, availability, and service health
- Support platform upgrades, maintenance, and operational readiness reviews
- Configure and maintain InfluxDB, including retention policies, performance tuning, and capacity planning
- Develop operational dashboards and reports for performance insights
- Acknowledge, investigate, troubleshoot, and resolve incidents
- Coordinate incident resolution with Infrastructure, Network, Cloud, Security, Application, and Service Delivery teams
- Participate in major incident bridges, disaster recovery exercises, and 24x7 operations support
- Follow escalation procedures, SOPs, runbooks, and ITIL processes; support Problem Management and RCA documentation
- Administer IBM Instana environments and support APM configuration, alerts, baselines, and thresholds
- Develop Python, PowerShell, Bash, and VBScript automation for operational tasks, monitoring deployments, remediation, and event-driven workflows
- Implement Ansible and Puppet infrastructure automation, server provisioning, configuration management, agent deployment, playbooks, and pipelines
- Maintain SOPs, runbooks, monitoring procedures, escalation matrices, and observability documentation
- Participate in knowledge-transfer sessions, service onboarding, service transition, migration, and continuous improvement initiatives
Requirements
What you’ll need- Required skills in Grafana, IBM Instana, SolarWinds, Telegraf, Prometheus, InfluxDB, OpenTelemetry, Grafana Alloy, APM monitoring, event management, alert management, observability concepts, and SLO/SLA monitoring
- FOAK support and Enterprise Logging & Telemetry (ELT) experience
- Log aggregation and correlation, telemetry data analysis, event correlation, application onboarding, and monitoring standards/observability frameworks
- VMware, Linux administration, Windows Server, Citrix VDI, DNS services, proxy services, middleware technologies, and infrastructure performance monitoring
- Python, PowerShell, Linux shell scripting, and VBScript
- Ansible, Puppet, webhooks, and Infrastructure as Code (IaC)
- ServiceNow, incident management, problem management, change management, ITIL Framework, and major incident management
- Preferred experience: 7 to 10+ years in Monitoring, Observability, Infrastructure Operations, SRE, or Platform Engineering
- Experience supporting large-scale enterprise environments and 24x7 operations
- Hands-on experience with Grafana, Instana, SolarWinds, Telegraf, Prometheus, and InfluxDB
- Experience supporting VMware, Citrix, middleware, enterprise applications, and cloud monitoring platforms
- Experience with FOAK applications and ELT platforms
- Strong troubleshooting, RCA, incident management, and operational support skills
- Experience with automation frameworks and Infrastructure as Code (Ansible preferred)
- Experience integrating observability platforms with enterprise automation solutions
- Nice-to-have: Docker, Kubernetes, AWS, Microsoft Azure, Google Cloud Platform, Jenkins, GitHub Actions, GitLab CI/CD, REST APIs, microservices monitoring, and DevOps/SRE practices
Benefits
Comp & perks- Unlimited Paid Days Off
- Three health plan options
- 401k with company match
- Dental, vision, short-term disability, long-term disability, life and AD&D coverage
- Flexible spending accounts
- Family Forming Benefit including fertility coverage and adoption/surrogacy reimbursement
- Paid childbearing and paternal leave
- Education Reimbursement
- Student Loan Assistance or 529 College Funding
- Sabbatical leave
- Wellness program
- Flexible work schedule
- Annual bonus plan based on company and individual performance
- Equity grant under the Associate Equity Appreciation Program
- Option to work from home or in Ensono offices when not required on a client site
