Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Baseten

Software Engineer – Inference Platform

Baseten

. Build infrastructure and orchestration systems for large-scale distributed LLM inference, including routing, autoscaling, scheduling, and runtime management .

Posted 10/6/2026full-timeUnited StatesMid-LevelSenior💰 $180,000 - $360,000 per yearWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in building and operating large-scale distributed systems and backend infrastructure, with a strong focus on performance, reliability, and developer experience. Proficient in API design, versioning, and operational excellence, with a solid foundation in debugging and optimizing production systems.

Highest-signal resume keywords
Distributed Systems EngineeringAPI Design and ManagementKubernetes ExpertisePerformance OptimizationOperational Excellence

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Distributed SystemsBackend InfrastructureAPI VersioningPerformance DebuggingCapacity PlanningSLO ManagementRate LimitingAuthenticationAutoscalingService Orchestration
Soft Skills
Written CommunicationCollaborationDeveloper Experience Focus
Tools & Technologies
KubernetesLLM Inference EnginesObservability ToolingCI/CD SystemsRelease Automation
Industry Keywords
Large-Scale APIsInference EngineeringGPU WorkloadsOpen-Source InfrastructureMultimodal Serving

Tech Stack

Tools & technologies
Distributed SystemsKubernetes

About the role

Key responsibilities & impact
  • Build infrastructure and orchestration systems for large-scale distributed LLM inference, including routing, autoscaling, scheduling, and runtime management
  • Design, build, and operate Model APIs with structured outputs, tool/function calling, and multimodal serving
  • Implement API versioning, validation, usage metering, quotas, and authentication
  • Instrument metrics, traces, and logs and build repeatable benchmarks for speed, reliability, and quality
  • Establish best practices for testing, release automation, and operational excellence
  • Debug and harden production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads
  • Partner with Inference Performance engineers and other teams to deliver optimizations to customers
  • Own projects end to end from architecture through deployment, monitoring, and iteration based on customer feedback
  • Balance performance, reliability, operational simplicity, and developer experience

Requirements

What you’ll need
  • Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, or a related field, or equivalent practical experience
  • 3+ years building and operating distributed systems, backend infrastructure, or large-scale APIs where reliability, latency, and scale are first-class concerns
  • Proven track record of owning low-latency, reliable backend services, including rate limiting, auth, quotas, metering, and migrations
  • Infrastructure instincts with experience in profiling, tracing, capacity planning, and SLO management
  • Comfort debugging performance and reliability issues across application, runtime, and infrastructure layers
  • Strong developer-experience focus
  • Eagerness to learn new languages, frameworks, and systems, and interest in inference engineering
  • Excellent written communication and collaboration skills, including writing clear design documents and working across functions
  • Experience with or contributions to LLM inference engines and frameworks such as vLLM, SGLang, TensorRT-LLM, TGI, or Dynamo is a plus
  • Deep Kubernetes experience, including operators and custom resources, plus familiarity with service meshes or API gateways is a plus
  • Experience with distributed scheduling, autoscaling, or service orchestration is a plus
  • Experience operating GPU workloads in production is a plus
  • Background in developer-facing infrastructure or APIs, or contributions to open-source infrastructure or ML systems is a plus
  • Familiarity with observability tooling, CI/CD systems, or release automation is a plus

Benefits

Comp & perks
  • Competitive compensation, including meaningful equity
  • (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break
  • Paid parental leave
  • Fertility and family-building stipend through Carrot
  • (U.S. only) Company-facilitated 401(k)
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities