Machine Learning Lead

Process Nine TechnologiesIndia
Adzuna INPosted 58m agoOriginal Listing
it-jobs

Job Description

ML Leads JD Key Responsibilities - Model Training & Fine-Tuning: Build, fine-tune, and optimize state-of-the-art NLP, LLM, Speech, and Vision models for scheduled Indian languages, utilizing parameter-efficient methods (LoRA, QLoRA, PEFT). - Indic Tokenization & Linguistics: Architect custom tokenizers and text-normalization pipelines to address the "fertility problem" in Devanagari, Dravidian, and other regional scripts, ensuring low-latency and cost-effective model inference. - Multimodal System Design: Develop robust OCR engines capable of parsing complex script geometries (conjoint consonants, Shirorekha, vowel modifiers) and integrate them into document intelligence pipelines. - Speech Engineering: Deploy and scale robust STT (Speech-to-Text) and TTS (Text-to-Speech) pipelines capable of handling heavy code-mixing (e.g., Hinglish, Tanglish), regional accents, and localized dialects. - Vernacular Guardrails & Evaluation: Establish culturally contextual benchmark datasets and implement safety guardrails. - Production Deployment (MLOps): Package and serve models using high-throughput frameworks (vLLM, Triton, ONNX) optimized for GPU environments, minimizing computational overhead for massive cross-lingual workloads. - Vernacular Fraud & Anomaly Detection: Architect risk-scoring systems and anomaly detection models capable of identifying fraud patterns in native scripts and code-mixed formats. Essential Qualifications & Technical Skills - Education: Bachelor’s or Master's degree in Computer Science, Mathematics, Statistics, or a closely related quantitative field. - Experience: 4+ years of professional experience building and deploying machine learning models in production environments, with a proven track record in Indian Language NLP, Speech, or Anomaly Detection. - Programming: Expert-level proficiency in Python and standard ML frameworks ( PyTorch , TensorFlow). - Indic AI Stack: Direct, hands-on experience with specialized Indic frameworks and datasets (e.g., AI4Bharat's IndicTrans2/IndicWhisper , Bhashini API, Kathbath, Sarvam-105B, or Aksharantar). - Fraud Stack: Proficiency in tabular/graph-based ML toolkits (XGBoost, LightGBM, PyTorch Geometric) and handling highly imbalanced target variables (SMOTE, class weights). - NLP & LLMs: Deep understanding of Transformer architectures, sequence-to-sequence modeling, cross-lingual embeddings, vector databases (Milvus, Pinecone, Qdrant), and quantization tools (bitsandbytes, GPTQ). - Speech & Vision Processing: Experience processing raw audio signals (grapheme-to-phoneme conversion, spectrogram analysis) or document structures using OCR networks (CRAFT, DBNet, LayoutLM). - Handling Code-Mixing: Proven ability to build models that gracefully parse text or speech containing heavy code-switching (mixed Latin/regional scripts, multi-language grammar). Skills:- Machine Learning (ML), Natural Language Processing (NLP), TensorFlow, Deep Learning, PyTorch, Fine-tuning LLMs, LoRA / QLoRA, Speech-to-Text (STT), Text-to-Speech (TTS), OCR and MLOps

Get AI-Matched to This Job

Upload your resume and our AI will score how well you match this and thousands of similar roles.