FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and operating large-scale distributed systems, with a focus on reliability, scalability, and security. Proven ability to lead complex infrastructure projects and mentor engineering teams while driving alignment across multiple stakeholders.
Highest-signal resume keywords
Large-Scale Distributed SystemsCloud Infrastructure (AWS, GCP)KubernetesProgramming Languages (Python, Rust, Go, Java)Technical Leadership
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Infrastructure DesignSystem ReliabilityScalability OptimizationArchitectural Decision-MakingSoftware Engineering FundamentalsIncident ResponseInfrastructure-As-CodeMachine Learning InfrastructureLinux Kernel TuningEBPF
Soft Skills
MentoringCommunicationCollaboration
Certifications & Qualifications
Bachelor's Degree
Industry Keywords
Infrastructure ProjectsPerformance BottlenecksSecurity EngineeringPrivacy EngineeringPostmortemsOn-Call Rotations
Tech Stack
Tools & technologiesAWSCloudDistributed SystemsGoogle Cloud PlatformJavaKubernetesLinuxPythonRustGo
About the role
Key responsibilities & impact- Scope and lead complex, multi-month infrastructure projects from ambiguity through production
- Build reliable, scalable, and performant infrastructure for Research teams
- Develop partnerships with researchers and Research teams
- Resolve performance and scalability bottlenecks
- Mentor engineers and raise the team's technical bar
- Drive alignment on technical direction across multiple teams
- Own the reliability, scalability, and security of systems as usage and complexity grow
- Improve Infrastructure operational processes, including incident response, postmortems, and on-call rotations
Requirements
What you’ll need- Experience designing, building, and operating large-scale distributed systems or infrastructure in production
- Track record of independently scoping and delivering complex, ambiguous, multi-month technical projects
- Prior experience as a technical lead or mentor for other engineers
- Experience making architectural decisions that other engineers and teams build on
- Strong software engineering fundamentals
- Proficiency in at least one programming language, such as Python, Rust, Go, or Java
- Experience with modern cloud infrastructure, including Kubernetes and infrastructure-as-code, on AWS and/or GCP
- Strong written and verbal communication skills
- Experience driving alignment across multiple teams or stakeholders
- Bachelor's degree or an equivalent combination of education, training, and/or experience
- Relevant field of study demonstrated through coursework, training, or professional experience
- Preferred: 10+ years of software engineering experience, excluding internships
- Preferred: Experience with machine learning infrastructure, including GPUs, TPUs, or Trainium, and networking infrastructure such as NCCL
- Preferred: Low-level systems experience, such as Linux kernel tuning or eBPF
- Preferred: Background in security or privacy engineering best practices
Benefits
Comp & perks- Optional equity donation matching
- Generous vacation
- Parental leave
- Flexible working hours
- Office space for collaboration
- Visa sponsorship support
- Immigration lawyer assistance
