Well-funded AI startup building real-time generative experiences through breakthrough inference efficiency.
Kernel Engineer (ML Accelerators)
Your mornings start profiling a fused attention kernel that's bottlenecking token generation. By afternoon you're rewriting the Triton or CUDA implementation, shaving microseconds off latency for a model that serves millions of concurrent users. This is engine-room work: every cycle matters, every memory bandwidth decision ripples to the product surface. You'll collaborate with researchers who think in tensor diagrams and production engineers who count p99 latencies in their sleep.
What they're looking for
- 2+ years shipping high-performance compute kernels in CUDA, Triton, or similar DSLs for ML workloads
- Deep intuition for GPU microarchitecture—memory hierarchies, occupancy, warp scheduling, and instruction-level tradeoffs
- Fluency in C/C++ and Python; ability to move between bare-metal optimization and PyTorch frontend code
- Track record profiling and optimizing transformer-based models, especially attention mechanisms and MoE routing
- Experience translating research prototypes into production-grade kernels with robust numerics and backward compatibility