FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Staff Infrastructure Engineer
The Hartford. Leverage agentic AI to identify, assess, and resolve critical service issues and major incidents .
Tech Stack
Tools & technologiesITSMSplunkSwift
About the role
Key responsibilities & impact- Leverage agentic AI to identify, assess, and resolve critical service issues and major incidents
- Ensure swift detection and response to service issues
- Partner with cross-functional teams for service restoration
- Continuously improve the event management process
- Manage technical and executive communication for major incidents
- Perform alert triage, correlation, and initial impact assessment
- Support major incident detection and escalation by validating symptoms, confirming affected services, and engaging resolver teams
- Use standard operating procedures, runbooks, and decision frameworks to investigate, prioritize, and escalate events
- Maintain situational awareness during active events and document event patterns and operational observations
- Promote events to incidents when thresholds are met
- Partner with internal and external teams during restoration activities
- Identify opportunities to improve monitoring effectiveness, event quality, automation, and operational readiness
- Collaborate with infrastructure, application, reliability engineering, service desk, and vendor teams
- Identify recurring alert issues, document operational insights, and recommend enhancements to monitoring, runbooks, and automation
Requirements
What you’ll need- 8+ years' experience in monitoring and observability tools such as Splunk, Dynatrace, ITSI, Moogsoft, ThousandEyes, or similar
- Experience in AI-driven data analysis and problem solving
- AI tool proficiency including prompt engineering and interrogation skills
- Familiarity with ITSM and ticketing platforms, including incident creation, categorization, escalation, and documentation
- Understanding of event correlation, alert prioritization, service impact analysis, and basic automation/workflow enablement
- Strong analytical thinking and operational judgment
- Ability to work under pressure during high-impact events
- Must be authorized to work in the US without company sponsorship
- Must be able to support a 24x7 operational model, including possible shift coverage, off-hours support, and escalation activities
Benefits
Comp & perks- Hybrid work schedule
- Short-term or annual bonuses
- Long-term incentives
- On-the-spot recognition