Senior AI/ML Engineer (Voice Intelligence & Speech AI)

HomeJobs Senior AI/ML Engineer (Voice Intelligence & Speech AI)

Senior AI/ML Engineer (Voice Intelligence & Speech AI)

Role Overview

We are seeking a highly skilled Senior AI/ML Engineer with extensive experience in Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Natural Language Understanding (NLU). The ideal candidate will lead the development of our Voice AI framework, specifically focusing on building an agentic platform for intelligent customer interactions and call analysis. You will be responsible for the end-to-end lifecycle of speech models, from fine-tuning Large Language Models (LLMs) for conversational intent extraction to deploying real-time voice synthesis and transcription services in a production environment.

Key Responsibilities

● Model Development & Fine-tuning: Lead the selection and fine-tuning of Speech-to-Text and LLM architectures to achieve >95% precision in multi-lingual conversation analysis across 12+ Indian regional languages.

● End-to-End Voice Pipelines: Design and optimize real-time and batch processing workflows to extract structured parameters—such as call disposition, payment commitments (PTP), and sentiment—from complex financial collection and support calls.

● Agentic Orchestration: Develop and maintain a runtime agent core with sub-agents for transcription, intelligence extraction, and decision logging to automate loan processing and verification.

● Voice Synthesis & Personality: Implement sophisticated voice bot personalities using advanced TTS (e.g., ElevenLabs) to manage interruptions, language switching, and empathetic customer engagement.

● Performance Evaluation: Conduct rigorous evaluation of audio quality and transcription accuracy (Word Error Rate, F1-score) to ensure high-reliability performance across production batches.

● Integration & Deployment: Collaborate with engineering to integrate third-party ASR/TTS services and deploy high-performance REST APIs within AWS (EKS/ECS) environments.

Key Responsibilities

● Model Development & Fine-tuning: Lead the selection and fine-tuning of Speech-to-Text and LLM architectures to achieve >95% precision in multi-lingual conversation analysis across 12+ Indian regional languages.

● End-to-End Voice Pipelines: Design and optimize real-time and batch processing workflows to extract structured parameters—such as call disposition, payment commitments (PTP), and sentiment—from complex financial collection and support calls.

● Agentic Orchestration: Develop and maintain a runtime agent core with sub-agents for transcription, intelligence extraction, and decision logging to automate loan processing and verification.

● Voice Synthesis & Personality: Implement sophisticated voice bot personalities using advanced TTS (e.g., ElevenLabs) to manage interruptions, language switching, and empathetic customer engagement.

● Performance Evaluation: Conduct rigorous evaluation of audio quality and transcription accuracy (Word Error Rate, F1-score) to ensure high-reliability performance across production batches.

● Integration & Deployment: Collaborate with engineering to integrate third-party ASR/TTS services and deploy high-performance REST APIs within AWS (EKS/ECS) environments.

Skills Must Have
  • Bachelor’s or Master’s degree in Computer Science, AI, Mathematics, or a related field.
Recruitment

Let us help your Recruitment.