Building identity infrastructure for autonomous agents.
Machine Learning Engineer - Evals
As agents proliferate across enterprise systems, measuring their reliability and behavior becomes a critical bottleneck. You will architect the evaluation frameworks that define how autonomous systems perform, fail, and adapt in the wild. This means designing rigorous metrics, building data pipelines with Kafka, and running experiments in PyTorch to stress-test agent identity at scale. The work directly determines whether these systems remain chaotic or become trustworthy.
What they're looking for
- 5 to 15 years of software or machine learning engineering experience, with deep fluency in Python
- Proven track record designing and implementing ML evaluation frameworks or benchmarking systems
- Strong production experience with PyTorch for model experimentation and Kafka for streaming data pipelines
- Ability to commute and work on-site in New York five days a week
- Comfort operating in an early-stage, highly technical environment with minimal process