FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and maintaining production ML inference APIs, with a strong focus on performance-sensitive services and integration across various frameworks and infrastructures. Proficient in Python and experienced in building stable APIs with robust validation and observability practices.
Highest-signal resume keywords
Python ProgrammingC++ DevelopmentML Inference FrameworksAPI Design and ValidationLinux and Kubernetes
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Machine Learning InferenceTokenizationPrompt FormattingStreaming GenerationContinuous BatchingKV-Cache ManagementError HandlingObservabilityVersioning PracticesPerformance Optimization
Tools & Technologies
VLLMHugging FacePyTorchTriton Inference ServerTensorRT-LLMContainersCI/CDGRPCRESTCerebras Runtime
Industry Keywords
Distributed SystemsProduction SoftwareLarge Language ModelsData-Intensive ApplicationsModel-Specific Features
Tech Stack
Tools & technologiesCloudDistributed SystemsGRPCKubernetesLinuxPythonPyTorchC++Go
About the role
Key responsibilities & impact- Design, implement, and maintain production ML inference APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, and multimodal inputs
- Create consistent request and response semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends
- Integrate emerging foundation models, tokenizers, prompt formats, sampling methods, attention variants, multimodal inputs, and model-specific features
- Maintain inference-interface compatibility while designing Cerebras-specific extensions, versioning, deprecation, validation, and backward-compatibility practices
- Extend and integrate custom inference services with vLLM, PyTorch, Hugging Face libraries, AMD ROCm, and Cerebras runtime components
- Build control and data paths coordinating GPU prefill with Cerebras decode, including request routing, state transfer, error handling, retries, and lifecycle management
- Optimize streaming, time to first token, latency, throughput, batching, serialization, tokenization, scheduling, and component communication
- Build validation systems for tokenization, sampling, logits, generated outputs, precision changes, model upgrades, determinism, and backend compatibility
- Define service indicators and build structured logging, tracing, metrics, dashboards, health checks, and diagnostic tooling
- Create conformance tests, workload-replay tools, model-validation suites, performance benchmarks, integration tests, and release gates
- Build configuration, SDKs, documentation, examples, debugging tools, and self-service workflows
- Partner with compiler, runtime, kernel, cloud, product, and solutions teams to translate requirements into scalable serving capabilities
Requirements
What you’ll need- 5+ years of software engineering experience, including substantial individual-contributor ownership of production software or distributed systems
- Strong programming ability in Python
- Experience developing performance-sensitive or highly concurrent services in C++, Go, or a similar systems language
- Hands-on experience with a model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face Text Generation Inference, or an equivalent platform
- Understanding of modern LLM inference concepts, including tokenization, prompt formatting, sampling, streaming generation, continuous batching, KV-cache management, and model configuration
- Experience integrating software across service, framework, runtime, and infrastructure boundaries
- Experience building stable APIs with validation, error handling, observability, compatibility, and versioning practices
- Experience with Linux, containers, Kubernetes or comparable orchestration systems, CI/CD, and operating latency-sensitive services in production
- Ability to diagnose correctness, reliability, and performance issues across distributed serving systems
- Bachelor’s degree in computer science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience
- Preferred: experience with OpenAI-compatible, gRPC, REST, or streaming inference APIs
- Preferred: experience modifying or contributing to vLLM, SGLang, PyTorch, Hugging Face Transformers, Triton, TensorRT-LLM, or another open-source ML systems project
- Preferred: experience enabling transformer, Mixture-of-Experts, diffusion, embedding, reranking, or multimodal model architectures
- Preferred: understanding of model-specific tokenization, chat templates, generation configuration, logits processing, stopping criteria, tool calling, structured generation, and constrained decoding
- Preferred: experience with disaggregated prefill/decode architectures, KV-cache transfer, prefix caching, chunked prefill, memory-aware admission control, or request scheduling
- Preferred: experience designing, building, or operating production APIs and services for machine learning, large language models, or other data-intensive applications
Benefits
Comp & perks- Build a breakthrough AI platform beyond the constraints of the GPU
- Publish and open source cutting-edge AI research
- Work on one of the fastest AI supercomputers in the world
- Job stability with startup vitality
- Simple, non-corporate work culture that respects individual beliefs
- Continuous learning, growth and support
- Equal and diverse work environment
