FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing, building, and operating distributed systems and cloud services, with a strong focus on performance, reliability, and security. Proven ability to lead cross-functional collaboration and drive engineering improvements through automation and best practices.
Highest-signal resume keywords
Go Programming LanguagePython Programming LanguageKubernetesCloud Services on AWS, GCP, AzureDistributed Systems Design
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Software EngineeringBackend Services DevelopmentInfrastructure AutomationEvent-Driven ArchitecturesSystem DesignConcurrency ReasoningError HandlingTestingPerformance OptimizationProduction Reliability
Soft Skills
Cross-Functional CollaborationMentorshipConsensus Building
Tools & Technologies
Cloud Control-Plane SystemsContainer OrchestrationDurable Workflow SystemsIdentity and Access ManagementUsage Metering and Billing
Industry Keywords
AI CloudGPU InfrastructureHPC EnvironmentsLarge-Scale AI/ML TrainingProduction Systems
Tech Stack
Tools & technologiesAWSAzureCloudDistributed SystemsGoogle Cloud PlatformKubernetesPythonGo
About the role
Key responsibilities & impact- Design, build, and operate services, APIs, control planes, and platform capabilities powering Lambda’s AI cloud
- Solve distributed-systems problems involving state, consistency, concurrency, scheduling, failure recovery, and safe lifecycle management
- Own the full engineering lifecycle from problem framing and architecture through implementation, testing, rollout, observability, on-call, and continuous improvement
- Improve system availability, latency, throughput, efficiency, security, and operability as Lambda scales
- Convert incidents and near misses into durable engineering improvements, including automation, testing, guardrails, and backstops
- Collaborate with product, infrastructure, networking, storage, security, and SRE teams to resolve dependencies and deliver customer outcomes
- Use AI-assisted development tools while independently verifying correctness, security, and maintainability
- Contribute to technical standards, design and code reviews, and mentorship
- At Staff level, lead cross-team architecture and raise the organization’s technical capabilities
- Operate distributed systems powering Lambda’s GPU cloud, including compute control planes, managed Kubernetes, cloud APIs, identity and access, usage metering and billing, capacity and orchestration, reliability, and developer-facing infrastructure
Requirements
What you’ll need- 7 or more years of professional software engineering experience, or equivalent evidence of impact building production systems
- Depth in at least one general-purpose programming language; Lambda primarily works in Go and Python
- Ability to reason about concurrency, error handling, and testing
- Experience designing, building, and operating backend services, distributed systems, infrastructure, or platform capabilities at meaningful scale
- Practical understanding of system design, data models, APIs, failure modes, performance, and production reliability tradeoffs
- Track record of owning complex work through delivery and operation, including testing, staged rollout, monitoring, incident response, and root-cause improvement
- Proven track record of aligning cross-functional partners and gaining consensus around decisions and tradeoffs
- Experience building cloud services or platform infrastructure, or operating large-scale production systems on AWS, GCP, Azure, or a comparable cloud platform
- Experience with Kubernetes, container orchestration, schedulers, controllers, or cloud control-plane systems
- Experience in cloud infrastructure or platform domains such as compute, storage, networking, identity and access, developer platforms, usage metering and billing, databases, or fleet management
- Experience with infrastructure automation, durable workflow systems, event-driven architectures, or infrastructure as code
- Experience designing highly available, multi-region, or rapidly scaling distributed systems
- Familiarity with GPU infrastructure, HPC environments, or large-scale AI/ML training and inference workloads
- Must be able and willing to work onsite at the San Francisco office 4 days a week
- Must be legally authorized to work in the United States; visa sponsorship may be available
Benefits
Comp & perks- Generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan
- AI interview recording opt-out without impact on candidacy
