Scaling real-time generative AI through unprecedented computational efficiency.
Infrastructure Engineer — Managed Inference
Generative models are shackled by inference bottlenecks; serving them at latency demanded by real-time interaction requires fundamentally rethinking infrastructure from the silicon up. You will architect and operate the managed inference platform that translates raw model weight into live, high-throughput experiences. This means designing Kubernetes orchestration and cloud deployments that push GPU utilization to its absolute limit, ensuring the pipeline never stalls between user prompt and model output.
What they're looking for
- 5-10 years building and operating large-scale distributed systems on AWS or GCP.
- Deep operational expertise with Kubernetes, Docker, and Terraform in production ML environments.
- Proven ability to deploy and maintain PyTorch inference workloads at significant scale.
- Fluency in observability tooling such as Prometheus and Grafana to diagnose complex performance regressions.
- Willingness to work on-site in San Francisco.