FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Principal Software Engineer
Digital Turbine. Define the long-term vision, architectural direction, and multi-year technology roadmap for high-scale ML platforms, model deployment infrastructure, and generative AI systems .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in architecting and deploying high-scale ML platforms and generative AI systems, with a strong focus on MLOps best practices and operational reliability. Proven ability to lead cross-departmental collaboration and communicate complex technical concepts to diverse stakeholders.
Highest-signal resume keywords
Architecting High-Throughput ML ModelsExpertise in PyTorch and TensorFlowCloud-Native Infrastructure MasteryDistributed Systems OptimizationMLOps Best Practices Implementation
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Machine Learning EngineeringModel Deployment InfrastructureReal-Time InferenceFeature EngineeringSystem ArchitectureDistributed ComputingPython ProgrammingC++ ProgrammingRust ProgrammingData Structures
Soft Skills
Mentoring EngineersCross-Departmental CollaborationStrategic AlignmentArchitectural ConsensusCommunication with Stakeholders
Tools & Technologies
KubernetesAWS SageMakerGCP Vertex AIMLflowKubeflowDeepSpeedMegatron-LMRayCI/CD PipelinesFeature Stores
Industry Keywords
Generative AIModel GovernanceOperational Reliability MetricsFoundation ModelsLarge-Scale SystemsReal-Time Production EnvironmentsVector DatabasesNoSQLGPU ClustersTPUs
Tech Stack
Tools & technologiesApacheAWSCloudDistributed SystemsGoogle Cloud PlatformKubernetesNoSQLPythonPyTorchRayRustSparkTensorflowC++
About the role
Key responsibilities & impact- Define the long-term vision, architectural direction, and multi-year technology roadmap for high-scale ML platforms, model deployment infrastructure, and generative AI systems
- Solve novel engineering challenges in model scaling, distributed systems, real-time inference, and feature engineering
- Evaluate ML research, frameworks, hardware accelerators, and third-party platforms for strategic technology and build-vs-buy decisions
- Establish company-wide MLOps best practices, security standards, evaluation frameworks, model governance protocols, and operational reliability metrics
- Mentor and technically sponsor Staff, Senior, and peer engineers without direct administrative management responsibilities
- Lead cross-departmental Architecture Review Boards and approve critical system designs, data pipelines, and production deployment architectures
- Architect and optimize distributed training clusters, feature platforms, model serving engines, and inference pipelines
- Lead integration of foundation models, LLM orchestration, RAG, PEFT, and vector infrastructure
- Design observability systems for model drift, system health, data quality, security posture, and business impact
- Partner with VPs, Directors, Product Leaders, and Domain Experts to translate business vision into technical roadmaps and architectural specifications
- Communicate technical concepts, trade-offs, risks, and investments to executives and non-technical stakeholders
- Build collaboration across ML Engineering, Data Engineering, Infrastructure/DevOps, Enterprise Security, and Product teams
Requirements
What you’ll need- Bachelor’s degree in Computer Science, Machine Learning, Data Science, or a related quantitative field required; Master’s or Ph.D. preferred
- 10+ years of software/ML engineering experience with a Bachelor’s degree, OR 8+ years of experience with an advanced degree
- Proven track record operating as a Principal (P5) or Staff (P4) level engineer on enterprise-scale systems
- Proven track record architecting, deploying, and maintaining high-throughput, low-latency, mission-critical ML models and pipelines in real-time production environments
- Expert-level mastery of PyTorch, TensorFlow, Jax, and modern distributed ML training frameworks such as DeepSpeed, Megatron-LM, and Ray
- Comprehensive mastery of cloud-native infrastructure, Kubernetes, feature stores, CI/CD pipelines, AWS SageMaker, GCP Vertex AI, MLflow, and Kubeflow
- Exceptional proficiency in system architecture, distributed computing, memory optimization, data structures, and Python, C++, or Rust
- Demonstrated ability to drive strategic alignment, architectural consensus, and engineering compliance across non-reporting teams and business units
- Ability to balance immediate execution needs with long-term architectural stability, scalability, and cost efficiency
- Preferred expertise in large-scale Generative AI architectures, foundation model pre-training/fine-tuning, guardrailing, agentic workflows, and high-performance vector retrieval systems
- Preferred knowledge of Ray, Apache Spark, Dask, vector databases, NoSQL, and distributed caching
- Preferred experience with GPU clusters, TPUs, custom silicon, CUDA, TensorRT, ONNX Runtime, and vLLM
- Preferred active open-source contribution, conference speaking, or peer-reviewed publications in machine learning or distributed systems
- Candidates must be local to the posting location
Benefits
Comp & perks- Hybrid work environment
- Equal opportunity employer committed to diversity and inclusion
- Inclusive, equitable, and culturally fluent environment
- Employer awards and recognition as a workplace of choice