FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and scaling distributed Reinforcement Learning (RL) systems, with a strong focus on post-training LLMs and high-throughput rollout generation. Proficient in building scalable RL environments and reward infrastructures while ensuring training efficiency and stability across GPU clusters.
Highest-signal resume keywords
Post-Training LLMs ExperienceDistributed PyTorch TrainingRL Environment DevelopmentGPU Clusters UnderstandingContainerization and Orchestration
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Reinforcement Learning (PPO/GRPO-family, RLHF, RLVR)Distributed PyTorchRL Post-Training Frameworks (veRL, OpenRLHF, TRL, Ray)Rollout Inference Engines (vLLM, SGLang)Reward Function DevelopmentEvaluation HarnessesSandboxed ExecutionMulti-Turn Tool UseNCCLMPI
Tools & Technologies
KubernetesRay
Industry Keywords
Large-Scale RLGPU ClustersNetworkingCommunication LibrariesResearch Contributions in RL
Tech Stack
Tools & technologiesKubernetesPyTorchRay
About the role
Key responsibilities & impact- Design, build, and scale distributed RL post-training systems across thousands of GPUs
- Build high-throughput rollout generation integrating vLLM, SGLang, weight synchronization, and asynchronous/off-policy schemes
- Design scalable RL environments for agentic, multi-step tasks including sandboxed code execution, tool use, computer use, and multimodal interaction
- Build reward infrastructure including verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking
- Develop evaluation, monitoring, and debugging tooling for stable large-scale RL runs
- Advance training efficiency and stability and turn post-training ideas into production runs with researchers
- Learn the current RL stack, diagnose bottlenecks, ship and validate improvements, and harden the full loop across thousands of GPUs
Requirements
What you’ll need- Hands-on experience post-training LLMs with RL (PPO/GRPO-family, RLHF, RLVR) at meaningful scale
- Extensive distributed PyTorch training and parallelism (FSDP, Tensor/Pipeline/Expert Parallel) for foundation models
- Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents, including sandboxed execution and multi-turn tool use
- Deep familiarity with RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang)
- Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference workloads
- Containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads
- Research contributions in RL for LLMs, or open-source contributions to RL training frameworks
Benefits
Comp & perks- Equal opportunity employer
- Flexible remote work arrangement in the EU
