FREE ACCESS
5,000–10,000 jobs/day
See all jobs on Scoutfield
Search thousands of fresh jobs every day.
Discover
- Fresh listings
- Fast filters
- No subscription required
Create a free account and start exploring right away.
Core Competencies
Role fitCore Competencies
Use this summary to align your resume positioning with the role.
Demonstrates expertise in designing and scaling distributed Reinforcement Learning (RL) systems, with a strong focus on post-training LLMs and high-throughput rollout generation. Proficient in building RL environments and reward infrastructures while ensuring stability and efficiency across large GPU clusters.
Highest-signal resume keywords
Post-Training LLMs ExperienceDistributed PyTorch TrainingRL Environment DevelopmentGPU Clusters UnderstandingRL Post-Training Frameworks Familiarity
ATS Keywords
Tailor your resumeApplicant Tracking System Keywords
Tip: use these terms in your resume and cover letter to boost ATS matches.
Hard Skills
Reinforcement Learning (PPO/GRPO-family, RLHF, RLVR)Distributed PyTorchWeight SynchronizationReward Infrastructure DevelopmentEvaluation HarnessesMulti-Turn Tool UseAsynchronous ArchitecturesDebugging ToolingRollout Inference EnginesReward Functions
Tools & Technologies
VLLMSGLangNCCLMPIKubernetesRay
Industry Keywords
Large-Scale RLGPU ClustersMixed Training WorkloadsContainerizationOrchestration
Tech Stack
Tools & technologiesKubernetesPyTorchRay
About the role
Key responsibilities & impact- Design, build, and scale distributed RL post-training systems across thousands of GPUs
- Build high-throughput rollout generation integrating vLLM, SGLang, weight synchronization, and asynchronous/off-policy schemes
- Design RL environments for agentic, multi-step tasks including sandboxed code execution, tool use, computer use, and multimodal interaction
- Build reward infrastructure including verifiable/programmatic rewards, reward-model serving, LLM-as-judge pipelines, and defenses against reward hacking
- Develop evaluation, monitoring, and debugging tooling for stable large-scale RL runs
- Advance training efficiency and stability and turn post-training ideas into production runs with researchers
- Learn the current RL stack and diagnose throughput, stability, and correctness issues
- Improve and validate a component of the RL loop on a real run
- Harden the full loop across thousands of GPUs and asynchronous architectures
Requirements
What you’ll need- Hands-on experience post-training LLMs with RL (PPO/GRPO-family, RLHF, RLVR) at meaningful scale
- Extensive distributed PyTorch training and parallelism (FSDP, Tensor/Pipeline/Expert Parallel) for foundation models
- Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents, including sandboxed execution and multi-turn tool use
- Deep familiarity with RL post-training frameworks (veRL, OpenRLHF, TRL, Ray orchestration) and rollout inference engines (vLLM, SGLang)
- Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI) under mixed training and inference workloads
- Nice to have: running RL training across 100+ GPUs, including asynchronous or disaggregated trainer/rollout architectures
- Nice to have: containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads
- Nice to have: research contributions in RL for LLMs or open-source contributions to RL training frameworks
Benefits
Comp & perks- Equal opportunity employer
- Voluntary diversity and inclusion survey; participation does not affect the job application
