Early-stage startup building autonomous AI agents for GPU kernel optimization

Member of Technical Staff

San FranciscoOn-site$200K - $200K1 - 5 years
You'll spend your days crafting Python and C++ systems that let AI agents autonomously rewrite GPU kernels, squeezing milliseconds out of inference pipelines. Backend work meets machine learning here: expect to build AWS infrastructure that deploys these agents, profile their outputs against real workloads, and iterate on the optimization strategies they discover. Three engineers sit near you in a San Francisco office; decisions move fast, but the problems reward patience—this is low-level performance work with tangible, measurable wins.

What they're looking for

  • 1–5 years shipping production code in Python and C or C++, with at least one role touching systems-level or ML infrastructure work
  • Solid grasp of GPU architecture, CUDA programming, or kernel-level performance tuning
  • Experience deploying and monitoring services on AWS, including IAM, networking, and cost-aware design
  • Comfort working in-person alongside a technical founding team with minimal process overhead
  • Track record of measuring before optimizing—profiling tools, benchmarking rigor, and reproducible experiments

Tech stack

PythonAWSCC++

Apply for this role

We use AI to help match your experience to the roles where you fit best. A member of our team reviews every application, and all hiring decisions are made by people.

This role is one we're recruiting for on behalf of a client company; the client's identity is kept confidential at this stage. A Fluency recruiter will follow up with details.

Share this role