FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Tech Stack
Tools & technologiesApacheAWSJavaKafkaKubernetesPythonPyTorchSparkTensorflowTerraformUnity
About the role
Key responsibilities & impact- Set technical direction across ML training, serving, and observability
- Serve as final escalation point for complex infrastructure problems involving GPU capacity, Spark tuning, and production incidents
- Own and evolve the paved-road framework, including shared CI/CD, model-workflow scaffolding, and Databricks Asset Bundles
- Lead architecture for LLM endpoint serving across Databricks, AWS, Snowflake, and self-hosted deployments
- Address latency, cost, caching, evaluation, and PHI-safe routing for LLM serving
- Own standards and tooling for MLflow, model registry, training image supply chain, and observability
- Partner with Data & ML Platform, Data Science, App Dev, and Operations teams
- Provide technical input to vendor and platform selection decisions
- Mentor senior engineers and guide platform consumers
- Write high-leverage code and Infrastructure-as-Code hands-on
Requirements
What you’ll need- 10+ years of software engineering experience, including 3+ years designing, evolving, and operating enterprise-scale ML platforms in production
- Strong technical judgment under ambiguity and track record of setting standards and influencing peers
- Production experience with Databricks and/or Amazon SageMaker
- Experience with MLflow or equivalent tracking and registry system
- Experience with at least one core ML framework, such as PyTorch or TensorFlow
- Fluency in Java or a JVM equivalent and Python
- Deep Apache Spark experience for large-scale data and distributed compute
- Deep AWS experience, including networking, IAM, GPU compute, storage, and messaging services
- Fluency with Terraform, containers, Kubernetes, and GitHub-based CI/CD for ML workloads
- Direct experience serving LLMs in production, including cost management, evaluation harnesses, and safe handling of sensitive prompts and outputs
- Daily use of Claude Code, Cursor, Copilot, or equivalent AI coding tools
- Clear written and verbal communication, especially in asynchronous remote settings
- This job is not eligible for employment sponsorship
- Preferred: technical leadership on healthcare or regulated-industry ML platforms; Databricks Asset Bundles, Unity Catalog, Iceberg, Delta; specialized inference pipelines; GPU capacity planning; Kafka or Kinesis; clinical or safety-sensitive AI evaluation and red teaming; open-source ML infrastructure or production ML publications
Benefits
Comp & perks- Total rewards compensation strategy
- Post-offer health screenings and vaccinations as required by clients
- Reasonable accommodations for individuals with physical and mental disabilities
- Equal employment opportunity protections
