Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Peraton

Senior Kubernetes Engineer

Peraton

. Support development, modernization, and migration of containerized and agentic workloads in a government multi-cloud environment with 70+ customer tenants .

Posted 9/23/2026full-timeRemote • United StatesSenior💰 $104,000 - $166,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates extensive expertise in managing and operating Azure Kubernetes Service (AKS) and Agentic Runtime for Kubernetes (ARK) in a multi-tenant government environment, ensuring compliance with FedRAMP standards and implementing effective incident management strategies. Proficient in Kubernetes architecture, observability standards, and risk mitigation to optimize performance and security.

Highest-signal resume keywords
Azure Kubernetes Service (AKS)Agentic Runtime for Kubernetes (ARK)FedRAMP ComplianceIncident ManagementKubernetes Architecture

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
KubernetesCNI NetworkingCSI StorageRBACHelmTerraformBicepPythonPowerShellGo
Soft Skills
LeadershipCommunicationTechnical Documentation
Tools & Technologies
Azure MonitorManaged PrometheusContainer InsightsAzure Log AnalyticsOpenTelemetryGitOpsCI/CD PipelinesAzure DevOpsGitHub Actions
Certifications & Qualifications
CKACKADAZ-104AZ-305
Industry Keywords
Multi-Cloud EnvironmentGovernmentIncident ResponseChange ManagementObservability Standards

Tech Stack

Tools & technologies
AzureCloudFluxKubernetesNode.jsPrometheusPythonTerraformGo

About the role

Key responsibilities & impact
  • Support development, modernization, and migration of containerized and agentic workloads in a government multi-cloud environment with 70+ customer tenants
  • Design and lead implementation of container-platform requirements and architecture on Azure Kubernetes Service (AKS)
  • Stand up and operate Agentic Runtime for Kubernetes (ARK) to host AI-agent workloads
  • Serve as the Tier 3 escalation point for complex incidents and outages across AKS and the ARK agentic-runtime environment
  • Investigate and resolve issues across AKS clusters, control plane, CNI networking, ingress, CSI storage, workload scheduling, ARK control plane, and ARK Custom Resources
  • Drive root cause analysis on high-severity incidents and implement solutions to prevent recurrence
  • Manage high-priority incident communications with stakeholders, leadership, and affected tenants
  • Define and enforce observability standards using Azure Monitor, Managed Prometheus, Container Insights, Azure Log Analytics, and OpenTelemetry, including ARK agent-level telemetry
  • Identify and mitigate cluster, node-pool, and workload risks proactively
  • Analyze performance, cost, and capacity trends and architect optimizations for stability, efficiency, and security at scale
  • Manage and maintain AKS clusters and the ARK agentic runtime across a 70+ tenant environment
  • Deploy, upgrade, and operate ARK on Kubernetes, managing CRDs, controllers, and model and tool integrations including MCP servers
  • Ensure tenant resources are secure, FedRAMP compliant, and continuously optimized by enforcing Kubernetes RBAC, network policy, and Azure workload identity
  • Partner with the Architecture team on Azure Well-Architected Framework and Kubernetes best-practice solutions
  • Own and enforce Change Management procedures
  • Lead incident intake, triage prioritization, and team workload management
  • Mentor and train Tier 2 engineers and lead knowledge-transfer sessions
  • Improve team metrics and service delivery with the Service Delivery Manager and Product Owner
  • Maintain documentation of incident response procedures, architecture decisions, and lessons learned

Requirements

What you’ll need
  • Bachelor's degree and 8 years of experience, or a Master's degree and 6 years of experience, or an Associate's degree and 10 years of experience, or a High School diploma/equivalent and 12 years of experience
  • Must be a U.S. Citizen with the ability to obtain and maintain a DHS Public Trust clearance
  • Engineering and systems analysis experience with a deep hands-on focus on Kubernetes in production
  • Extensive hands-on expertise operating Azure Kubernetes Service (AKS) and the broader Kubernetes stack, including control plane, CNI networking, ingress, CSI storage, RBAC, Helm, and workload scheduling
  • Hands-on experience deploying and operating Agentic Runtime for Kubernetes (ARK) or comparable Kubernetes-native agent or operator platforms built on the CRD-plus-controller pattern, including ARK Custom Resources and MCP/A2A integration
  • Proven experience leading incident management and root cause analysis in large-scale, multi-tenant government environments
  • Deep knowledge of FedRAMP compliance and Kubernetes and Azure security best practices
  • Strong leadership, communication, and technical documentation skills
  • Ability to communicate effectively with technical teams and senior stakeholders
  • Preferred: CKA or CKAD, plus AZ-104 or AZ-305
  • Preferred: proficiency with GitOps and Infrastructure as Code, including Terraform, Bicep, Helm, Argo CD or Flux, and CI/CD pipelines using Azure DevOps or GitHub Actions
  • Preferred: advanced scripting with Python, PowerShell, or Go, including use of the ARK Python SDK
  • Preferred: experience integrating LLM providers such as Azure OpenAI, Anthropic, Google, or local models behind an agentic runtime