Apply

Ready to go for it?

AI Apply speeds things up—apply directly if you prefer.

FREE ACCESS
5,000–10,000 jobs/day
Scoutfield Logo

See all jobs on Scoutfield

Search thousands of fresh jobs every day.

Discover
  • Fresh listings
  • Fast filters
  • No subscription required
Create a free account and start exploring right away.
Anyone AI

GPU Kernel Engineer – CUDA, Triton, Accelerator Performance

Anyone AI

. Review GPU and accelerator kernel implementations for correctness .

Posted 9/15/2026contractRemote • ArgentinaMid-LevelSenior💰 $65 per hourWebsite

Core Competencies

Role fit
Core Competencies

Use this summary to align your resume positioning with the role.

Demonstrates expertise in developing, optimizing, and debugging GPU and accelerator kernels, with a strong focus on performance optimization and numerical correctness. Proficient in using profiling tools and translating kernels across frameworks while identifying and resolving technical challenges.

Highest-signal resume keywords
CUDATritonGPU Performance OptimizationKernel Profiling ToolsFloating-Point Numerical Correctness

ATS Keywords

Tailor your resume
Applicant Tracking System Keywords

Tip: use these terms in your resume and cover letter to boost ATS matches.

Hard Skills
GPU Kernel DevelopmentKernel OptimizationDebugging Kernel IssuesKernel ProfilingMemory Hierarchy OptimizationPerformance BenchmarkingKernel TranslationOperator FusionPerformance Bottleneck IdentificationNumerical Tolerance Evaluation
Soft Skills
Clear Technical FeedbackProblem-Solving
Tools & Technologies
NsightNCURoofline AnalysisCuBLASCuDNNAWS NeuronTPUJAXMLIRXLA
Industry Keywords
GPU EcosystemAccelerator EcosystemAI Workload PerformanceKernel LibrariesTechnical Benchmark Development

Tech Stack

Tools & technologies
AWS

About the role

Key responsibilities & impact
  • Review GPU and accelerator kernel implementations for correctness
  • Compare outputs against reference implementations
  • Evaluate numerical tolerance thresholds
  • Review kernel benchmarks and determine whether comparisons are fair
  • Identify performance bottlenecks and optimization opportunities
  • Assess whether performance targets are realistic given hardware limits
  • Review kernel translations and hardware migrations
  • Identify compilation, driver, memory, shape, and runtime issues
  • Determine whether technical tasks are genuinely difficult or incorrectly configured
  • Provide clear, actionable technical feedback
  • Implement and debug kernels
  • Optimize CUDA and Triton kernels
  • Translate between kernel frameworks
  • Perform hardware migration and operator fusion
  • Profile and benchmark performance
  • Verify numerical correctness
  • Debug compilation and runtime issues
  • Optimize memory hierarchy and kernel-level AI workload performance

Requirements

What you’ll need
  • 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels
  • Strong experience with at least two of: CUDA; Triton; NKI / AWS Neuron; Pallas / JAX
  • Strong understanding of GPU performance optimization
  • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers
  • Understanding of memory bandwidth, compute throughput, GPU occupancy, shared memory, register pressure, memory coalescing, and bank conflicts
  • Strong understanding of floating-point numerical correctness and tolerance thresholds
  • Experience debugging kernel compilation and runtime issues
  • Ability to distinguish software defects, environment problems, and genuine optimization challenges
  • Experience writing kernels from technical specifications, translating kernels between frameworks, migrating kernels across hardware platforms, debugging incorrect implementations, optimizing kernel performance, and fusing multiple operations into optimized kernels
  • Experience across both NVIDIA GPU and custom accelerator ecosystems (nice to have)
  • Experience with AWS Trainium, TPU, JAX, or other accelerators (nice to have)
  • Compiler engineering experience (nice to have)
  • Familiarity with MLIR, XLA, or intermediate representation lowering (nice to have)
  • Contributions to GPU or ML kernel libraries (nice to have)
  • Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls (nice to have)
  • Experience with AI model evaluation, RLHF, or technical benchmark development (nice to have)

Benefits

Comp & perks
  • $65 per hour compensation
  • Part-time, project-based consulting engagement
  • Remote work