Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
fal

Senior Software Engineer, Machine Learning Infrastructure – Automation

fal

. Own, design, build, and maintain ML CI/CD infrastructure for automated testing, validation, and deployment pipelines for ML models and inference services .

Posted 10/9/2026full-timeSan Francisco • California • United StatesSenior💰 $170,000 - $230,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and maintaining ML CI/CD infrastructure, with a strong focus on automation, performance testing, and developer productivity. Proficient in Python and experienced in designing reliable distributed systems and optimizing CI/CD pipelines at scale.

Highest-signal resume keywords
Python ProficiencyCI/CD System DesignAutomated TestingML Infrastructure ExperienceDocker and Cloud Infrastructure

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
CI/CD InfrastructureAutomated TestingPerformance TestingDistributed Systems DesignDependency ManagementParallel ExecutionML Quality EvaluationAgentic Coding ToolsInfrastructure-as-CodeNVIDIA GPU Architectures
Soft Skills
Problem IdentificationIndependent WorkCollaboration
Tools & Technologies
GitHub ActionsDockerCloud InfrastructurePyTorchCodexClaude Code
Industry Keywords
ML ModelsInference ServicesAutomated ValidationDeveloper ProductivityCompute Resources

Tech Stack

Tools & technologies
CloudDistributed SystemsDockerPythonPyTorch

About the role

Key responsibilities & impact
  • Own, design, build, and maintain ML CI/CD infrastructure for automated testing, validation, and deployment pipelines for ML models and inference services
  • Reduce CI execution times through intelligent parallelization, caching, test selection, and efficient use of compute resources
  • Develop automated systems to test model outputs, detect quality regressions, and validate changes across models, GPU architectures, and configurations
  • Build continuous performance testing for inference latency, throughput, GPU utilization, and cost
  • Automate validation of model pricing, billing configurations, API schemas, and deployments before production
  • Develop AI-powered automation and agentic coding systems to diagnose CI failures, identify regressions, propose fixes, and streamline engineering workflows
  • Build deployment safeguards, verification, rollback mechanisms, and monitoring
  • Identify repetitive ML-team tasks and build automation tools to eliminate engineering toil
  • Collaborate closely with Applied ML and ML Performance teams
  • Build infrastructure, tools, and model access that help generative media teams move from idea to production at scale

Requirements

What you’ll need
  • 5+ years of software engineering background with proficiency in Python and experience building production infrastructure and developer tooling
  • Experience designing and operating CI/CD systems using GitHub Actions or comparable technologies
  • Deep understanding of automated testing, build systems, dependency management, caching, and parallel execution
  • Experience working with containerized workloads, Docker, and cloud infrastructure
  • Ability to design reliable distributed systems and debug complex infrastructure failures
  • Strong understanding of observability, including logs, metrics, tracing, and automated alerting
  • A passion for developer productivity and a demonstrated ability to eliminate manual processes through automation
  • Comfortable working independently, identifying high-impact problems, and building end-to-end solutions
  • Experience with ML infrastructure, PyTorch, GPU workloads, or model-serving systems
  • Familiarity with NVIDIA GPU architectures and multi-GPU environments
  • Experience building automated inference benchmarks or ML quality evaluation frameworks
  • Experience with agentic coding tools such as Codex or Claude Code, or building custom AI engineering agents
  • Experience optimizing CI/CD pipelines at scale, including distributed test execution and ephemeral compute environments
  • Experience developing internal developer platforms or infrastructure-as-code tooling
  • Must be legally authorized to work in the United States
  • Must answer whether visa sponsorship is required to work in the United States
  • Must be able to meet the requirement for in-person attendance five days per week in downtown San Francisco

Benefits

Comp & perks
  • Offers equity
  • Interesting and challenging work
  • A lot of learning and growth opportunities
  • Health, dental, and vision insurance (US)
  • Regular team events and offsites
  • Access to fal's massive GPU cluster for testing and development