Diffusion-based LLM startup making AI radically faster and cheaper.

Machine Learning Systems Engineer

South Bay AreaOn-site$200K - $300K2 - 5 years
You'll spend your days squeezing inference latency out of diffusion models—rewriting kernels, profiling memory bandwidth, and deciding when a custom CUDA op beats PyTorch's defaults. One morning might mean benchmarking a new attention variant on H100s; the afternoon, debugging why a quantized checkpoint OOMs at batch size eight. This is work where shaving milliseconds translates directly to dollars saved for customers running million-token generations.

What they're looking for

  • 2–5 years shipping ML infrastructure in production, with solid footing in Python and at least one systems language (C/C++, Rust, or Go)
  • Hands-on experience optimizing deep learning training or inference—kernel fusion, custom operators, memory planning, or distributed strategies you've actually debugged
  • Comfortable navigating cloud GPU stacks: you've set up clusters on AWS, GCP, or Azure, and have opinions on Kubernetes schedulers for ML workloads
  • Track record with at least one major framework (PyTorch or TensorFlow) including lower-level APIs like torch.fx, torch.compile, or XLA
  • Genuine curiosity about diffusion architectures—you've read the papers, maybe reimplemented a sampler, and get why consistency models matter for throughput

Tech stack

PythonCC++RustGoPyTorchTensorFlowDocker

Apply for this role

We use AI to help match your experience to the roles where you fit best. A member of our team reviews every application, and all hiring decisions are made by people.

This role is one we're recruiting for on behalf of a client company; the client's identity is kept confidential at this stage. A Fluency recruiter will follow up with details.

Share this role