Building voice AI that handles millions of customer interactions with human-level understanding.
Site Reliability Engineer
Referral bonus eligible. Know someone who would be great for this role? Refer them to a Fluency Digital recruiter. If they're placed, you may be eligible for a referral bonus of up to $5,000 (terms apply). Refer someone for this role
You'll shape how a growing AI platform stays dependable at scale—designing resilient infrastructure, taming complex incident response, and making sure the systems that power real-time conversational agents never miss a beat. This team treats reliability as a feature, not an afterthought, blending software engineering with deep operational discipline to protect customer-facing services. Expect to write code that improves observability, harden Kubernetes workloads on GCP, and partner closely with engineers who ship TypeScript, NodeJS, and Python into production daily. If tracing a latency spike through a mesh of microservices sounds like a good Tuesday, you'll fit right in.
What they're looking for
- 5–8 years in SRE, DevOps, or infrastructure engineering roles, with a track record of running production systems at scale
- Deep fluency with Kubernetes and GCP, including cluster management, networking, and cost-aware architecture
- Strong coding ability in Python or TypeScript/NodeJS for automation, tooling, and reliability improvements
- Experience owning observability stacks (Prometheus, Grafana, and related alerting pipelines) and driving incident postmortems that lead to real change
- Comfort designing and evolving CI/CD pipelines alongside Terraform-managed infrastructure