Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Cerebras

Machine Learning Engineer – Model Bring-Up

Cerebras

. Bring up new models by understanding architectures, loading and converting weights, implementing supported execution paths, and establishing correctness against reference implementations .

Posted 10/9/2026full-timeBengaluru • IndiaMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in optimizing machine learning models through advanced programming in C++ and Python, with a strong focus on MLIR and performance tuning for AI accelerators. Proficient in debugging and profiling to ensure model correctness and efficiency across various execution environments.

Highest-signal resume keywords
C++ ProgrammingPython ProgrammingMLIR ExperienceModel OptimizationGPU Profiling

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Machine Learning Model DebuggingTransformer ArchitecturesQuantization TechniquesCompiler FundamentalsTensor OperationsAttention MechanismsPerformance BenchmarkingKernel DevelopmentDataflow AnalysisIntermediate Representations
Tools & Technologies
PyTorchLLVMAI AcceleratorsDistributed ExecutionModel Parallelism
Industry Keywords
MLIR DialectsOperator FusionKV CachingSpeculative DecodingLow-Bit Quantization

Tech Stack

Tools & technologies
PythonPyTorchC++

About the role

Key responsibilities & impact
  • Bring up new models by understanding architectures, loading and converting weights, implementing supported execution paths, and establishing correctness against reference implementations
  • Lower models to hardware using MLIR by developing and extending dialects, graph transformations, lowering passes, and hardware-specific mappings
  • Enable attention, matrix multiplication, normalization, positional embeddings, and other operators through compiler and kernel changes
  • Optimize operator fusion, tensor layouts, tiling, memory allocation, data movement, and parallel execution
  • Tune prefill and decode performance, KV cache management, batching, and quantization to improve latency, throughput, and memory efficiency
  • Investigate numerical differences and measure accuracy impact of precision changes and compiler optimizations
  • Use profiling, execution traces, and hardware counters to identify compute, memory, communication, and runtime limitations
  • Collaborate with hardware, compiler, kernel, and runtime teams to deliver reliable model support and repeatable performance benchmarks

Requirements

What you’ll need
  • Strong programming skills in C++ and Python
  • Hands-on experience bringing up and debugging ML models in PyTorch or a comparable framework
  • Practical experience with MLIR, including dialects, rewrite patterns, transformation passes, and lowering pipelines
  • Understanding of compiler fundamentals, including intermediate representations, dataflow analysis, and code generation
  • Understanding of transformer architectures, attention mechanisms, tensor operations, and numerical precision
  • Experience profiling and optimizing workloads on GPUs or other AI accelerators
  • Ability to debug correctness and performance issues across model code, compiler-generated code, kernels, and runtime execution
  • Experience with LLM inference, including GQA, sliding-window attention, MoE, KV caching, and speculative decoding
  • Experience with FP16, BF16, FP8, or low-bit quantization and their accuracy and performance tradeoffs
  • Experience developing accelerator kernels or hardware-specific compiler backends
  • Familiarity with distributed execution, model parallelism, and accelerator memory hierarchies
  • Contributions to MLIR, LLVM, inference frameworks, or related open-source projects

Benefits

Comp & perks
  • Job stability with startup vitality
  • Opportunity to publish and open source cutting-edge AI research
  • Work on one of the fastest AI supercomputers in the world
  • Simple, non-corporate work culture that respects individual beliefs
  • Continuous learning, growth and support
  • Equal and diverse work environment