Well-funded AI startup building real-time generative experiences through breakthrough inference efficiency.

Kernel Engineer (ML Accelerators)

San FranciscoOn-site$200K - $240K2+ years
Your mornings start profiling a fused attention kernel that's bottlenecking token generation. By afternoon you're rewriting the Triton or CUDA implementation, shaving microseconds off latency for a model that serves millions of concurrent users. This is engine-room work: every cycle matters, every memory bandwidth decision ripples to the product surface. You'll collaborate with researchers who think in tensor diagrams and production engineers who count p99 latencies in their sleep.

What they're looking for

  • 2+ years shipping high-performance compute kernels in CUDA, Triton, or similar DSLs for ML workloads
  • Deep intuition for GPU microarchitecture—memory hierarchies, occupancy, warp scheduling, and instruction-level tradeoffs
  • Fluency in C/C++ and Python; ability to move between bare-metal optimization and PyTorch frontend code
  • Track record profiling and optimizing transformer-based models, especially attention mechanisms and MoE routing
  • Experience translating research prototypes into production-grade kernels with robust numerics and backward compatibility

Tech stack

CC++PythonPyTorch

Apply for this role

We use AI to help match your experience to the roles where you fit best. A member of our team reviews every application, and all hiring decisions are made by people.

This role is one we're recruiting for on behalf of a client company; the client's identity is kept confidential at this stage. A Fluency recruiter will follow up with details.

Share this role