FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Principal Software Engineer
Digital Turbine. Define the long-term vision, architectural direction, and multi-year technology roadmap for high-scale ML platforms, model deployment infrastructure, and generative AI systems .
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in architecting and optimizing high-scale ML platforms, model deployment infrastructure, and generative AI systems. Proficient in establishing MLOps best practices and leading cross-departmental collaborations to translate business vision into technical roadmaps.
Highest-signal resume keywords
Architecting High-Throughput ML ModelsExpertise in PyTorch and TensorFlowCloud-Native Infrastructure MasteryDistributed Systems OptimizationMLOps Best Practices Implementation
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Machine Learning EngineeringModel Deployment InfrastructureReal-Time InferenceFeature EngineeringSystem ArchitectureDistributed ComputingPython ProgrammingC++ ProgrammingRust ProgrammingData Structures
Soft Skills
Mentoring EngineersTechnical SponsorshipCross-Departmental CollaborationCommunication with StakeholdersStrategic Alignment
Tools & Technologies
KubernetesAWS SageMakerGCP Vertex AIMLflowKubeflowDeepSpeedMegatron-LMRayVector DBsNoSQL
Certifications & Qualifications
Bachelor’s Degree in Computer ScienceMaster’s Degree in Machine LearningPh.D. in Data Science
Industry Keywords
Generative AIMLOpsModel GovernanceOperational ReliabilityHigh-Performance Vector Retrieval
Tech Stack
Tools & technologiesApacheAWSCloudDistributed SystemsGoogle Cloud PlatformKubernetesNoSQLPythonPyTorchRayRustSparkTensorflowC++
About the role
Key responsibilities & impact- Define the long-term vision, architectural direction, and multi-year technology roadmap for high-scale ML platforms, model deployment infrastructure, and generative AI systems
- Solve novel engineering challenges in model scaling, distributed systems, real-time inference, and feature engineering
- Evaluate frontier ML research, frameworks, hardware accelerators, and third-party vendor platforms for build-vs-buy decisions
- Establish company-wide MLOps best practices, security standards, evaluation frameworks, model governance protocols, and operational reliability metrics
- Mentor and technically sponsor Staff, Senior, and peer engineers without direct administrative management responsibilities
- Lead cross-departmental Architecture Review Boards and approve critical system designs, data pipelines, and production deployment architectures
- Foster technical excellence, continuous learning, operational resilience, and principled engineering tradeoffs
- Architect and optimize distributed training clusters, feature platforms, model serving engines, and inference pipelines
- Lead design and integration of foundation models, LLM orchestration, RAG, PEFT, and vector infrastructure
- Design observability systems for model drift, system health, data quality, security posture, and business metric impact
- Partner with VPs, Directors, Product Leaders, and Domain Experts to translate business vision into technical roadmaps and architectural specifications
- Communicate technical concepts, trade-offs, risks, and investments to executive leadership and non-technical stakeholders
- Collaborate across ML Engineering, Data Engineering, Infrastructure/DevOps, Enterprise Security, and Product teams
Requirements
What you’ll need- Bachelor’s degree in Computer Science, Machine Learning, Data Science, or a related quantitative field required; Master’s or Ph.D. preferred
- 10+ years of software/ML engineering experience with a Bachelor’s degree, OR 8+ years of experience with an advanced degree (Master’s / Ph.D.)
- Proven track record operating as a Principal (P5) or Staff (P4) level engineer on enterprise-scale systems
- Proven track record architecting, deploying, and maintaining high-throughput, low-latency, mission-critical ML models and pipelines in real-time production environments
- Expert-level mastery of PyTorch, TensorFlow, Jax, and modern distributed ML training frameworks such as DeepSpeed, Megatron-LM, and Ray
- Comprehensive mastery of cloud-native infrastructure, Kubernetes, feature stores, CI/CD pipelines, AWS SageMaker, GCP Vertex AI, MLflow, and Kubeflow
- Exceptional proficiency in system architecture, distributed computing, memory optimization, data structures, and Python, C++, or Rust
- Demonstrated ability to drive strategic alignment, architectural consensus, and engineering compliance across non-reporting teams and business units
- Ability to balance immediate execution needs with long-term architectural stability, scalability, and cost efficiency
- Preferred: expertise in large-scale Generative AI architectures, foundation model pre-training/fine-tuning, guardrailing, agentic workflows, and high-performance vector retrieval systems
- Preferred: knowledge of Ray, Apache Spark, Dask, Vector DBs, NoSQL, and distributed caching
- Preferred: experience with GPU clusters, TPUs, custom silicon, CUDA, TensorRT, ONNX Runtime, and vLLM
- Preferred: active open-source contribution, conference speaking, or peer-reviewed publications in machine learning or distributed systems
Benefits
Comp & perks- Hybrid work environment
- Equal opportunity employer committed to diversity and inclusion
- Inclusive, equitable, and culturally fluent environment