Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Cerebras

Staff Software Engineer, Inference API

Cerebras

. Design, implement, and maintain production ML inference APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, and multimodal inputs .

Posted 9/22/2026full-timeToronto • CanadaLeadWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and maintaining production ML inference APIs, with a strong focus on performance-sensitive services and integration across various frameworks and infrastructures. Proficient in Python and experienced in building stable APIs with robust validation and observability practices.

Highest-signal resume keywords
Python ProgrammingC++ DevelopmentML Inference FrameworksAPI Design and ValidationLinux and Kubernetes

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Machine Learning InferenceTokenizationPrompt FormattingStreaming GenerationContinuous BatchingKV-Cache ManagementError HandlingObservabilityVersioning PracticesPerformance Optimization
Tools & Technologies
VLLMHugging FacePyTorchTriton Inference ServerTensorRT-LLMContainersCI/CDGRPCRESTCerebras Runtime
Industry Keywords
Distributed SystemsProduction SoftwareLarge Language ModelsData-Intensive ApplicationsModel-Specific Features

Tech Stack

Tools & technologies
CloudDistributed SystemsGRPCKubernetesLinuxPythonPyTorchC++Go

About the role

Key responsibilities & impact
  • Design, implement, and maintain production ML inference APIs for chat completions, text generation, streaming, model configuration, tool calling, structured outputs, and multimodal inputs
  • Create consistent request and response semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends
  • Integrate emerging foundation models, tokenizers, prompt formats, sampling methods, attention variants, multimodal inputs, and model-specific features
  • Maintain inference-interface compatibility while designing Cerebras-specific extensions, versioning, deprecation, validation, and backward-compatibility practices
  • Extend and integrate custom inference services with vLLM, PyTorch, Hugging Face libraries, AMD ROCm, and Cerebras runtime components
  • Build control and data paths coordinating GPU prefill with Cerebras decode, including request routing, state transfer, error handling, retries, and lifecycle management
  • Optimize streaming, time to first token, latency, throughput, batching, serialization, tokenization, scheduling, and component communication
  • Build validation systems for tokenization, sampling, logits, generated outputs, precision changes, model upgrades, determinism, and backend compatibility
  • Define service indicators and build structured logging, tracing, metrics, dashboards, health checks, and diagnostic tooling
  • Create conformance tests, workload-replay tools, model-validation suites, performance benchmarks, integration tests, and release gates
  • Build configuration, SDKs, documentation, examples, debugging tools, and self-service workflows
  • Partner with compiler, runtime, kernel, cloud, product, and solutions teams to translate requirements into scalable serving capabilities

Requirements

What you’ll need
  • 5+ years of software engineering experience, including substantial individual-contributor ownership of production software or distributed systems
  • Strong programming ability in Python
  • Experience developing performance-sensitive or highly concurrent services in C++, Go, or a similar systems language
  • Hands-on experience with a model-serving framework such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face Text Generation Inference, or an equivalent platform
  • Understanding of modern LLM inference concepts, including tokenization, prompt formatting, sampling, streaming generation, continuous batching, KV-cache management, and model configuration
  • Experience integrating software across service, framework, runtime, and infrastructure boundaries
  • Experience building stable APIs with validation, error handling, observability, compatibility, and versioning practices
  • Experience with Linux, containers, Kubernetes or comparable orchestration systems, CI/CD, and operating latency-sensitive services in production
  • Ability to diagnose correctness, reliability, and performance issues across distributed serving systems
  • Bachelor’s degree in computer science, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent practical experience
  • Preferred: experience with OpenAI-compatible, gRPC, REST, or streaming inference APIs
  • Preferred: experience modifying or contributing to vLLM, SGLang, PyTorch, Hugging Face Transformers, Triton, TensorRT-LLM, or another open-source ML systems project
  • Preferred: experience enabling transformer, Mixture-of-Experts, diffusion, embedding, reranking, or multimodal model architectures
  • Preferred: understanding of model-specific tokenization, chat templates, generation configuration, logits processing, stopping criteria, tool calling, structured generation, and constrained decoding
  • Preferred: experience with disaggregated prefill/decode architectures, KV-cache transfer, prefix caching, chunked prefill, memory-aware admission control, or request scheduling
  • Preferred: experience designing, building, or operating production APIs and services for machine learning, large language models, or other data-intensive applications

Benefits

Comp & perks
  • Build a breakthrough AI platform beyond the constraints of the GPU
  • Publish and open source cutting-edge AI research
  • Work on one of the fastest AI supercomputers in the world
  • Job stability with startup vitality
  • Simple, non-corporate work culture that respects individual beliefs
  • Continuous learning, growth and support
  • Equal and diverse work environment