Research Engineer - Agent Intelligence & Evaluation

ixigoNew Delhi, Delhi
Adzuna INPosted 26m agoOriginal Listing
it-jobs

Job Description

Job Description Voice agents fail in ways traditional software doesn't. ASR confidence drops on an accent and a tool call misfires. Latency breaks turn-taking and the LLM hallucinates a policy. A model swap silently regresses production and nobody catches it for a week. We're building self-healing voice agents for enterprise customer support. This role owns the intelligence layer: the evals that catch failures before shipping, the observability that traces them across the pipeline, and the feedback loops that let agents fix themselves What you'll own • Evaluation infrastructure. Audio-native metrics for barge-in, prosody, and turn-taking. Adversarial datasets across accents and edge cases. LLM-as-judge rubrics for task success, tool-use correctness, and recovery. • Observability across the pipeline. Tracing that correlates audio, STT, LLM reasoning, tool calls, and TTS to a single conversation. Analysis and alerting that surfaces cascade failures instead of hiding them. • Self-improvement systems. Mine production traces for failure patterns, generate targeted training or prompt data, validate fixes with adversarial replay, and guardrail against regressions.

Get AI-Matched to This Job

Upload your resume and our AI will score how well you match this and thousands of similar roles.