Research Fellowship (Applied AI/ML)

ixigoNew Delhi, Delhi
Adzuna INPosted 22h agoOriginal Listing
it-jobs

Job Description

Job Description Research Fellowship: Agent Intelligence & Evaluation Voice agents fail in ways traditional software doesn't. An ASR confidence drop on a regional accent misfires a tool call, an LLM hallucinates a policy because upstream latency broke turn-taking, and support teams roll these agents back within a week without anyone able to explain what went wrong. What you'll work on Over 4 months, you'll take on one or two of the following, shaped by your interests. Evaluation frameworks. Text-only evals miss most of what matters in voice: barge-in, prosody, latency-induced errors, cross-turn context loss. You'll design audio-native metrics, generate adversarial conversational datasets across accents and edge cases, and build LLM-as-judge rubrics for task completion, empathy, and recovery from tool failures. End-to-end observability. Tracing a failed interaction means correlating audio packets, STT hypotheses, LLM reasoning traces, tool calls, and TTS output back to a single conversation ID. You'll help shape the schema and analysis layer that makes cascade failures visible across the stack. Self-improvement systems. Once you can measure and trace, the interesting work is closing the loop: mining production traces for failure patterns, generating targeted fine-tuning data or prompt updates, and validating that fixes hold under adversarial replay. Who we're looking for Someone who cares about the research questions for their own sake, and equally cares whether the work ships. Papers at Interspeech, ACL, NeurIPS, or EMNLP on speech, dialogue systems, agent evaluation, or human-AI interaction are directly relevant. Comfortable in Python, and familiar with at least one of: speech models (Whisper, Conformer variants), LLM tool-use and agent frameworks, or observability stacks (OpenTelemetry, Langfuse, Arize, Hamming). Current PhD students in ML, NLP, or speech are the strong default; exceptional MS students or research engineers with a publication track record are welcome to apply. Nice to have Prior work on evaluation methodology, dataset synthesis, or interpretability. Experience with real-time systems, telephony, or streaming pipelines. A blog, repo, or workshop paper that shows how you think in public. What you'll get ₹50,000/month for the 4-month term, access to real enterprise conversation data under proper governance, mentorship on the research and shipping sides, co-authorship on papers that come out of the work, and a system in production running on top of what you build.

Get AI-Matched to This Job

Upload your resume and our AI will score how well you match this and thousands of similar roles.