Preference Model
RL Environments Engineer - Low-Level Engineering and Kernel Inference Optimization
Remote Senior $4.7k–$12.8k/moest.
Summary
We are hiring a Low-Level Engineer to design and build RL environments that teach LLMs kernel development, hardware optimization, and systems programming. The role focuses on creating realistic feedback loops where models learn to write high-performance code across GPU and CPU architectures.
What you'll do
- Design and develop RL environments for kernel optimization and systems programming
- Build feedback mechanisms for teaching LLMs high-performance code generation
- Optimize code execution across GPU and CPU architectures
- Develop and iterate on training environments for hardware-aware learning
- Debug and optimize kernel performance using CUDA and HIP/ROCm
Requirements
- Strong engineering-quality Python skills
- Production mindset with focus on debugging, reliability, and iteration speed
- Clear understanding of LLMs and their current limitations
- Deep knowledge of memory hierarchies (registers, L1/L2/shared memory, HBM, system RAM)
- Experience with threading models, synchronization primitives, and concurrent programming
- Knowledge of cache coherence, memory access patterns, coalescing, and bank conflicts
- Experience with JIT frameworks (Triton, JAX/XLA, TorchInductor, Numba) and AOT compilation (LLVM, MLIR, TVM)
- Proficiency in modern C++ including templates, concurrency, and build systems
- Assembly-level programming and low-level optimization experience (x86, ARM, NVIDIA Hopper/Blackwell)
- CUDA and/or HIP/ROCm kernel optimization expertise
- PyTorch custom operators and backend extension development
- GPU communication libraries experience (NCCL, RCCL, MPI, UCX)
- Mixed-precision and low-precision kernel optimization knowledge
- Advanced English proficiency (C1/C2)
Conditions
Schedule: Remote contractor role with 4 hours overlap to PST
Location: Remote
Language: Advanced English required (C1/C2)