Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
CACI International Inc

SRE Platform Engineer

CACI International Inc

. Monitor, maintain, and support production and non-production environments for OIGChat, AI applications, and enterprise data platforms .

Posted 9/23/2026full-timeUnited StatesSeniorLead💰 $114,600 - $252,100 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive expertise in Site Reliability Engineering (SRE) principles, including monitoring, observability, incident response, and performance optimization. Proficient in managing Azure cloud services and operating AI/ML platforms in production environments.

Highest-signal resume keywords
Site Reliability Engineering (SRE)Azure Cloud ServicesAzure MonitorIncident ResponseAI/ML Platforms

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
MonitoringObservabilityCapacity PlanningPerformance OptimizationTroubleshootingBackup/Restore ProcessesDisaster RecoveryHigh AvailabilityLog AnalysisSynthetic Monitoring
Soft Skills
Problem SolvingLeadershipCollaboration
Tools & Technologies
Azure DatabricksGrafanaPrometheusELK StackApplication Insights
Certifications & Qualifications
Active DHS/EOD Clearance
Industry Keywords
Federal GovernmentMission-Critical Systems24/7 Operational EnvironmentsPlatform EngineeringDevOps

Tech Stack

Tools & technologies
AWSAzureCloudGrafanaPrometheus

About the role

Key responsibilities & impact
  • Monitor, maintain, and support production and non-production environments for OIGChat, AI applications, and enterprise data platforms
  • Implement observability, alerting, dashboards, health checks, synthetic monitoring, and log analysis
  • Lead incident response, troubleshooting, root cause analysis, defect resolution, dependency updates, and emergency changes
  • Analyze performance and usage metrics across applications, APIs, AI model endpoints, data pipelines, and infrastructure
  • Support capacity planning, resource sizing, autoscaling configuration, and cost optimization
  • Implement and maintain backup/restore processes, disaster recovery, high availability, and business continuity capabilities
  • Monitor Azure Databricks clusters, data pipelines, storage services, and analytical workloads
  • Support deployment and operational readiness for pilot applications and new capabilities
  • Provide surge support for complex technical issues, large-scale data collection analysis, environment optimization, and specialized troubleshooting
  • Develop and maintain runbooks, troubleshooting guides, architecture diagrams, incident post-mortems, and knowledge transfer materials

Requirements

What you’ll need
  • Bachelor's degree + 15 years of experience in site reliability engineering, DevOps, platform engineering, systems administration, or related field; equivalencies considered (Master's + 12 years; 21 years with no degree; AA + 17 years)
  • Must be able to obtain an Active DHS/EOD Clearance as required
  • Extensive experience with Site Reliability Engineering (SRE) principles including monitoring, observability, incident response, capacity planning, performance optimization, and reliability engineering practices
  • Proven expertise with Azure cloud services including compute, storage, networking, monitoring, and platform-as-a-service (PaaS) offerings
  • Strong experience with Azure Monitor, Application Insights, Grafana, Prometheus, and ELK stack
  • Demonstrated ability to troubleshoot complex technical issues across application, platform, and infrastructure layers
  • Experience operating AI/ML platforms, Azure OpenAI, Databricks, Synapse, or high-scale cloud applications in production environments
  • Hands-on experience with Azure Government or other secure government cloud environments such as AWS GovCloud
  • Background in federal government, mission-critical systems, or 24/7 operational environments

Benefits

Comp & perks
  • Flexible time off benefit
  • Robust learning resources
  • Healthcare benefits
  • Wellness benefits
  • Financial benefits
  • Retirement benefits
  • Family support benefits
  • Continuing education benefits
  • Time off benefits
  • Competitive compensation
  • Benefits and learning and development opportunities