South Bay AI lab building foundation models for real-world deployment
Member of Technical Staff, Inference Systems
Referral bonus eligible. Know someone who would be great for this role? Refer them to a Fluency Digital recruiter. If they're placed, you may be eligible for a referral bonus of up to $5,000 (terms apply). Refer someone for this role
You'll spend your days squeezing milliseconds from transformer inference—profiling kernels, rewriting attention blocks in Rust, and pushing batch throughput until the hardware sweats. The work sits at the collision of research engineering and systems programming: your Python experiments become C++ that ships. Expect to own a slice of the serving stack end-to-end, from tensor-memory layouts to the RPC layer hitting the accelerator.
What they're looking for
- 2+ years shipping high-performance systems in Rust, C/C++, or Go, with at least one stint close to ML workloads
- Solid grasp of GPU architectures and CUDA/ROCm—enough to read a kernel profiler and know where the warps stall
- Comfortable dropping into PyTorch internals to patch or extend, then rewriting the hot path in a systems language
- Experience with distributed serving at scale: batching strategies, speculative decoding, or pipeline parallelism
- Track record of measurable latency or throughput wins on production inference, not just benchmark leaderboard entries