FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.

Senior Software Engineer, Machine Learning Infrastructure – Automation
fal. Own, design, build, and maintain ML CI/CD infrastructure for automated testing, validation, and deployment pipelines for ML models and inference services .
Posted 10/9/2026full-timeSan Francisco • California • United StatesSenior💰 $170,000 - $230,000 per yearWebsite
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in building and maintaining ML CI/CD infrastructure, with a strong focus on automation, performance testing, and developer productivity. Proficient in Python and experienced in designing reliable distributed systems and optimizing CI/CD pipelines at scale.
Highest-signal resume keywords
Python ProficiencyCI/CD System DesignAutomated TestingML Infrastructure ExperienceDocker and Cloud Infrastructure
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
CI/CD InfrastructureAutomated TestingPerformance TestingDistributed Systems DesignDependency ManagementParallel ExecutionML Quality EvaluationAgentic Coding ToolsInfrastructure-as-CodeNVIDIA GPU Architectures
Soft Skills
Problem IdentificationIndependent WorkCollaboration
Tools & Technologies
GitHub ActionsDockerCloud InfrastructurePyTorchCodexClaude Code
Industry Keywords
ML ModelsInference ServicesAutomated ValidationDeveloper ProductivityCompute Resources
Tech Stack
Tools & technologiesCloudDistributed SystemsDockerPythonPyTorch
About the role
Key responsibilities & impact- Own, design, build, and maintain ML CI/CD infrastructure for automated testing, validation, and deployment pipelines for ML models and inference services
- Reduce CI execution times through intelligent parallelization, caching, test selection, and efficient use of compute resources
- Develop automated systems to test model outputs, detect quality regressions, and validate changes across models, GPU architectures, and configurations
- Build continuous performance testing for inference latency, throughput, GPU utilization, and cost
- Automate validation of model pricing, billing configurations, API schemas, and deployments before production
- Develop AI-powered automation and agentic coding systems to diagnose CI failures, identify regressions, propose fixes, and streamline engineering workflows
- Build deployment safeguards, verification, rollback mechanisms, and monitoring
- Identify repetitive ML-team tasks and build automation tools to eliminate engineering toil
- Collaborate closely with Applied ML and ML Performance teams
- Build infrastructure, tools, and model access that help generative media teams move from idea to production at scale
Requirements
What you’ll need- 5+ years of software engineering background with proficiency in Python and experience building production infrastructure and developer tooling
- Experience designing and operating CI/CD systems using GitHub Actions or comparable technologies
- Deep understanding of automated testing, build systems, dependency management, caching, and parallel execution
- Experience working with containerized workloads, Docker, and cloud infrastructure
- Ability to design reliable distributed systems and debug complex infrastructure failures
- Strong understanding of observability, including logs, metrics, tracing, and automated alerting
- A passion for developer productivity and a demonstrated ability to eliminate manual processes through automation
- Comfortable working independently, identifying high-impact problems, and building end-to-end solutions
- Experience with ML infrastructure, PyTorch, GPU workloads, or model-serving systems
- Familiarity with NVIDIA GPU architectures and multi-GPU environments
- Experience building automated inference benchmarks or ML quality evaluation frameworks
- Experience with agentic coding tools such as Codex or Claude Code, or building custom AI engineering agents
- Experience optimizing CI/CD pipelines at scale, including distributed test execution and ephemeral compute environments
- Experience developing internal developer platforms or infrastructure-as-code tooling
- Must be legally authorized to work in the United States
- Must answer whether visa sponsorship is required to work in the United States
- Must be able to meet the requirement for in-person attendance five days per week in downtown San Francisco
Benefits
Comp & perks- Offers equity
- Interesting and challenging work
- A lot of learning and growth opportunities
- Health, dental, and vision insurance (US)
- Regular team events and offsites
- Access to fal's massive GPU cluster for testing and development