Early-stage AI infrastructure company building tools for multimodal model deployment
Member of Technical Staff, ML Systems
You'll spend your days tracing memory allocations through distributed PyTorch training runs, identifying where gradients balloon or checkpointing falls through. The work sits at the intersection of compiler-level optimization and systems engineering — squeezing performance out of hardware configurations that change month to month as new accelerators ship. Your benchmarks directly shape how robotics and media customers ship perception models to production.
What they're looking for
- 1+ years shipping production ML systems or distributed training infrastructure
- Deep familiarity with PyTorch internals: autograd, distributed data parallel, memory pooling
- Experience profiling and optimizing GPU/TPU utilization across multi-node clusters
- Solid systems fundamentals — you think in terms of thread scheduling, page faults, and PCIe bandwidth
- Comfort working on-site in the South Bay, embedded in a tight-knit technical team