Early-stage startup building autonomous AI agents for GPU kernel optimization
Member of Technical Staff
You'll spend your days crafting Python and C++ systems that let AI agents autonomously rewrite GPU kernels, squeezing milliseconds out of inference pipelines. Backend work meets machine learning here: expect to build AWS infrastructure that deploys these agents, profile their outputs against real workloads, and iterate on the optimization strategies they discover. Three engineers sit near you in a San Francisco office; decisions move fast, but the problems reward patience—this is low-level performance work with tangible, measurable wins.
What they're looking for
- 1–5 years shipping production code in Python and C or C++, with at least one role touching systems-level or ML infrastructure work
- Solid grasp of GPU architecture, CUDA programming, or kernel-level performance tuning
- Experience deploying and monitoring services on AWS, including IAM, networking, and cost-aware design
- Comfort working in-person alongside a technical founding team with minimal process overhead
- Track record of measuring before optimizing—profiling tools, benchmarking rigor, and reproducible experiments