Building identity infrastructure for autonomous agents.

Machine Learning Engineer - Evals

New YorkOn-site$220K - $300K5 - 15 years
As agents proliferate across enterprise systems, measuring their reliability and behavior becomes a critical bottleneck. You will architect the evaluation frameworks that define how autonomous systems perform, fail, and adapt in the wild. This means designing rigorous metrics, building data pipelines with Kafka, and running experiments in PyTorch to stress-test agent identity at scale. The work directly determines whether these systems remain chaotic or become trustworthy.

What they're looking for

  • 5 to 15 years of software or machine learning engineering experience, with deep fluency in Python
  • Proven track record designing and implementing ML evaluation frameworks or benchmarking systems
  • Strong production experience with PyTorch for model experimentation and Kafka for streaming data pipelines
  • Ability to commute and work on-site in New York five days a week
  • Comfort operating in an early-stage, highly technical environment with minimal process

Tech stack

PyTorchKafkaPython

Apply for this role

We use AI to help match your experience to the roles where you fit best. A member of our team reviews every application, and all hiring decisions are made by people.

This role is one we're recruiting for on behalf of a client company; the client's identity is kept confidential at this stage. A Fluency recruiter will follow up with details.

Share this role