FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and implementing complex, distributed systems with a strong focus on AI/ML integration, observability, and operational readiness. Proven ability to mentor engineering teams and establish best practices for software development and automation.
Highest-signal resume keywords
Distributed Systems ArchitectureAI/ML IntegrationREST API DesignCI/CD and Automated TestingObservability and Monitoring
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
PythonJavaJavaScriptPowerShellSQLNoSQLEvent-Driven SystemsData PipelinesRisk ScoringAnomaly Detection
Soft Skills
MentoringCollaborationInfluencing Engineering Decisions
Tools & Technologies
AzureGCPServiceNowITSM Platforms
Industry Keywords
Change ManagementOperational ReadinessData QualitySecurity and GovernanceModel Evaluation
Tech Stack
Tools & technologiesAzureCloudDistributed SystemsGoogle Cloud PlatformITSMJavaJavaScriptNoSQLPythonServiceNowSQL
About the role
Key responsibilities & impact- Lead architecture and detailed design for complex, distributed, and highly integrated systems
- Establish engineering standards, reusable patterns, and best practices across software, APIs, automation, and AI/ML capabilities
- Lead technical design reviews and influence engineering decisions across teams and organizations
- Identify opportunities to simplify, modernize, and standardize the technology landscape
- Mentor engineers and raise the technical capabilities of the broader engineering organization
- Build intelligent workflows for change request creation, risk assessment, approval routing, scheduling, validation, and verification
- Develop services and automation that improve change success, failure, and health metrics through instrumentation and data-driven engineering
- Apply modular architecture, automated testing, observability, security, and operational readiness
- Design event-driven and API-based architectures for reliable, scalable system-to-system integration
- Build automation for pre-change validation, dependency and impact analysis, risk assessment, and post-change verification
- Maintain data consistency, integrity, lineage, and traceability across enterprise systems
- Enable ingestion and retrieval of topology, configuration, policy, incident, and operational data
- Develop reusable integration patterns and platform capabilities
- Architect and integrate AI/ML and LLM-based capabilities into enterprise engineering workflows
- Apply anomaly detection, classification, recommendation systems, risk scoring, and LLM-based reasoning
- Implement retrieval, grounding, context management, and evaluation strategies for LLM applications
- Establish guardrails, observability, evaluation, and fallback mechanisms for production AI systems
- Partner with engineering, product, architecture, and operations stakeholders to move AI capabilities into enterprise-scale production
- Design for high availability, fault tolerance, scalability, and graceful degradation
- Establish SLOs, SLIs, observability, alerting, and operational practices for critical services
- Identify and resolve performance bottlenecks across applications, APIs, data stores, and distributed systems
- Drive improvements based on production telemetry, incidents, and operational trends
- Ensure failure handling, resiliency, recovery, and rollback strategies
Requirements
What you’ll need- Bachelor's degree in computer science, computer engineering, computer information systems, software engineering, or related area and 4 years’ experience in software engineering or related area; OR 6 years’ experience in software engineering or related area
- Extensive experience building production-grade distributed systems
- Strong programming experience in Python, Java, JavaScript, PowerShell, or equivalent languages
- Experience designing REST APIs, event-driven systems, webhooks, and asynchronous workflows
- Strong understanding of CI/CD, automated testing, and operational readiness
- Hands-on experience with Azure, GCP, or comparable cloud platforms
- Experience with SQL/NoSQL databases, data pipelines, caching, messaging/event platforms, and distributed data processing
- Strong understanding of observability, monitoring, logging, tracing, and production operations
- Experience with ServiceNow Change Management or comparable ITSM platforms
- Strong focus on data quality, consistency, lineage, security, and governance
- Experience integrating AI/ML models and/or LLM APIs into production applications
- Strong understanding of LLM application architecture, including model integration, context management, grounding/retrieval, evaluation, and guardrails
- Experience with risk scoring, anomaly detection, classification, recommendation systems, or other decision-support applications
- Understanding of model evaluation, monitoring, explainability, and continuous improvement
- Experience designing reliable and scalable AI-powered production systems, including latency, cost, accuracy, and failure handling considerations
Benefits
Comp & perks- Incentive awards for performance
- 401(k) match
- Stock purchase plan
- Paid maternity and parental leave
- PTO and/or PPTO, including sick leave, vacation, and holidays
- Multiple health plans, including medical, vision, and dental coverage
- Company-paid life insurance
- Family care leave
- Bereavement leave
- Jury duty and voting leave
- Short-term and long-term disability
- Company discounts
- Military Leave Pay
- Adoption and surrogacy expense reimbursement
- Live Better U Walmart-paid education benefit, including tuition, books, and fees
- Performance-based bonus awards
- Stock compensation may be included for certain positions
