Back to all testimonials
Industry
B2B Software & IT
Retail & Marketing Tech
Services

AI DevOps & Infrastructure

Read the case study

OpenCall: Real-time, low latency AI voice at enterprise scale

"OpenCall.ai powers AI-driven voice interactions at enterprise scale, where every millisecond of latency and every nuance of natural speech directly impacts the end-user experience. Tech 42 helped us take our fine-tuned Sesame CSM text-to-speech model from concept to production on AWS — delivering a dedicated, real-time voice generation layer with streaming audio and word-level alignment that our platform depends on.

The architecture is built for performance and reliability: GPU-backed ECS infrastructure, containerized model deployment via ECR, S3-based model loading, WebSocket audio streaming, CSM-native alignment metadata, API key authentication, and fully Terraform-managed infrastructure — all working together to give our platform precise, real-time visibility into what words are being spoken as audio is produced.

One of the most valuable contributions was Tech 42's approach to intelligent autoscaling — balancing GPU utilization, active Gateway connections, and TTS capacity signals to keep the service responsive without letting idle GPU capacity drive up costs. For a production speech model running at scale, that balance is everything.

Tech 42 also set our team up for long-term ownership with smoke-test scripts, operational documentation, model refresh guidance, API key rotation procedures, and a hands-on knowledge transfer session. Their depth in MLOps, AWS infrastructure, and real-time AI systems made them a true partner in extending our Sesame CSM platform — and the result is a voice experience that is more natural, more scalable, and fully production-ready.

What stood out beyond the technical delivery was the team itself. Tech 42 consistently went above and beyond — proactively solving problems, communicating with clarity, and showing genuine investment in our success. They felt less like a consulting partner and more like an extension of our own team. We look forward to continuing to partner with Tech 42 on future initiatives as we scale our platform."

‍

Arthur Silverstein

CTO & Co-founder

Project summary

OpenCall's voice AI platform for healthcare required sub-300ms latency to first audio chunk and independence from third-party speech vendors. Tech 42 containerized the Sesame CSM text-to-speech model and OpenAI's Whisper speech-to-text model, stored both in Amazon ECR, and deployed them on GPU-backed AWS ECS. An Application Load Balancer distributes TTS traffic across container instances. CloudWatch alarms drive ECS auto-scaling policies, allowing the infrastructure to handle variable concurrent call volumes without manual intervention. The deployed pipeline achieved approximately 200ms latency to first audio chunk, meeting and exceeding the 300ms production threshold. OpenCall took full ownership of the AWS-native voice stack, including operational documentation and smoke-test scripts.

Back to all testimonials