Scaling real-time generative AI through unprecedented computational efficiency.

Infrastructure Engineer — Managed Inference

San FranciscoOn-site$200K - $260K5 - 10 years
Generative models are shackled by inference bottlenecks; serving them at latency demanded by real-time interaction requires fundamentally rethinking infrastructure from the silicon up. You will architect and operate the managed inference platform that translates raw model weight into live, high-throughput experiences. This means designing Kubernetes orchestration and cloud deployments that push GPU utilization to its absolute limit, ensuring the pipeline never stalls between user prompt and model output.

What they're looking for

  • 5-10 years building and operating large-scale distributed systems on AWS or GCP.
  • Deep operational expertise with Kubernetes, Docker, and Terraform in production ML environments.
  • Proven ability to deploy and maintain PyTorch inference workloads at significant scale.
  • Fluency in observability tooling such as Prometheus and Grafana to diagnose complex performance regressions.
  • Willingness to work on-site in San Francisco.

Tech stack

KubernetesAWSGCPK8sPyTorchDockerTerraformPrometheus

Apply for this role

We use AI to help match your experience to the roles where you fit best. A member of our team reviews every application, and all hiring decisions are made by people.

This role is one we're recruiting for on behalf of a client company; the client's identity is kept confidential at this stage. A Fluency recruiter will follow up with details.

Share this role