Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Luma AI

Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma AI

. Design, build, and scale distributed RL post-training systems across thousands of GPUs .

Posted 10/8/2026full-timeRemoteMid-LevelSeniorWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in designing and scaling distributed Reinforcement Learning (RL) systems, with a strong focus on post-training LLMs and high-throughput rollout generation. Proficient in building scalable RL environments and reward infrastructures while ensuring training efficiency and stability across GPU clusters.

Highest-signal resume keywords
Post-Training LLMs ExperienceDistributed PyTorch TrainingRL Environment DevelopmentGPU Clusters UnderstandingContainerization and Orchestration

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
Reinforcement Learning (PPO/GRPO-family, RLHF, RLVR)Distributed PyTorchRL Post-Training Frameworks (veRL, OpenRLHF, TRL, Ray)Rollout Inference Engines (vLLM, SGLang)Reward Function DevelopmentEvaluation HarnessesSandboxed ExecutionMulti-Turn Tool UseNCCLMPI
Tools & Technologies
KubernetesRay
Industry Keywords
Large-Scale RLGPU ClustersNetworkingCommunication LibrariesResearch Contributions in RL

Tech Stack

Tools & technologies
KubernetesPyTorchRay

About the role

Key responsibilities & impact
  • Design, build, and scale distributed RL post-training systems across thousands of GPUs
  • Build high-throughput rollout generation integrating vLLM, SGLang, weight synchronization, and asynchronous/off-policy schemes
  • Design scalable RL environments for agentic, multi-step tasks including sandboxed code execution, tool use, computer use, and multimodal interaction
  • Build reward infrastructure including verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking
  • Develop evaluation, monitoring, and debugging tooling for stable large-scale RL runs
  • Advance training efficiency and stability and turn post-training ideas into production runs with researchers
  • Learn the current RL stack, diagnose bottlenecks, ship and validate improvements, and harden the full loop across thousands of GPUs

Requirements

What you’ll need
  • Hands-on experience post-training LLMs with RL (PPO/GRPO-family, RLHF, RLVR) at meaningful scale
  • Extensive distributed PyTorch training and parallelism (FSDP, Tensor/Pipeline/Expert Parallel) for foundation models
  • Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents, including sandboxed execution and multi-turn tool use
  • Deep familiarity with RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang)
  • Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference workloads
  • Containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads
  • Research contributions in RL for LLMs, or open-source contributions to RL training frameworks

Benefits

Comp & perks
  • Equal opportunity employer
  • Flexible remote work arrangement in the EU