FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in Python programming and cloud infrastructure management, particularly on Google Cloud Platform and Kubernetes, while applying machine learning fundamentals to optimize system design and performance for large-scale applications.
Highest-signal resume keywords
Production-Ready Python ProgrammingGoogle Cloud Platform ExperienceKubernetes ManagementMachine Learning FundamentalsSystem Design for High-Availability Infrastructure
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Python ProgrammingAlgorithmsData StructuresCloud InfrastructureKubernetesModel DeploymentPerformance TuningObservability ToolingTransformer ArchitecturesLLM Serving Infrastructure
Soft Skills
CollaborationCommunicationProblem-Solving
Tools & Technologies
Google Cloud PlatformKubernetesObservability Tooling
Industry Keywords
Machine LearningInfrastructure ManagementHigh-Availability SystemsOpen-Source ModelsProduction Workloads
Tech Stack
Tools & technologiesCloudGoogle Cloud PlatformKubernetesPython
About the role
Key responsibilities & impact- Write high-quality, production-grade Python code
- Participate in code reviews and pair programming
- Contribute to and help drive architectural decisions
- Build, operate, and improve infrastructure for training and serving machine learning models on Google Cloud Platform and Kubernetes
- Design and maintain serving infrastructure for open-source and internally hosted LLMs
- Manage model deployment and resources, and tune performance
- Apply machine learning fundamentals to infrastructure and design decisions
- Partner with Applied Scientists to understand model training and serving needs
- Translate model requirements into infrastructure that reduces workflow friction
- Contribute to system design discussions, balancing performance, cost, and reliability
- Support infrastructure serving more than 100 million users
- Use generative AI and productivity tools thoughtfully
Requirements
What you’ll need- Bachelor's degree in Computer Science, Applied Statistics, Mathematics, Electrical Engineering, or a related quantitative field, or equivalent professional experience
- 5+ years of professional experience building, iterating on, and troubleshooting complex backend and infrastructure systems
- Strong software engineering fundamentals, including algorithms and data structures
- Production-ready Python programming
- Hands-on experience with cloud infrastructure, preferably Google Cloud, and Kubernetes
- Experience deploying and operating production workloads
- Familiarity with observability tooling, including metrics, logging, and tracing
- Working knowledge of machine learning fundamentals and concepts
- Awareness of LLM serving infrastructure and hosting open-source models
- Basic understanding of transformer architectures, predictors, and how model design choices affect serving infrastructure
- Comfort with system design for large-scale, high-availability infrastructure
Benefits
Comp & perks- Equity package
- Annual performance bonus
- Competitive benefits supporting employees and their families
- Visa sponsorship may be available for certain roles and skills
- Hybrid/remote flexibility depending on proximity to the Brooklyn office
- Occasional travel arrangement for remote candidates outside commuting distance
