Sr Video ML Engineer (Founding Team)
Cliply Pte LtdIndia₹4,000,000 – ₹5,500,000
it-jobs
Job Description
Role Overview The Senior AI/ Machine Learning Engineer will design, build, and optimise the core intelligence layer that powers Cliply’s video understanding and content analysis platform. This role blends deep hands-on engineering with architectural ownership. You will work across video, audio, and text modalities, shaping model design while also delivering production-ready ML systems. You will be the technical anchor for Cliply’s AI stack, partnering closely with the Lead Architect and the engineering team to bring research concepts into scalable, real-world systems. Key Responsibilities Multimodal & Video ML Architecture - Design and validate deep learning architectures for video, audio, and text understanding, including temporal modelling and multimodal fusion. - Define approaches for long-sequence modelling, representation learning, and sequence-to-sequence tasks. - Lead experiments with transformers, vision transformers, video encoders, and hybrid multimodal architectures. Model Development & Optimisation - Build and optimise models for content understanding, highlight detection, ranking, and scoring. - Implement training pipelines, data loaders, augmentations, and evaluation metrics for large-scale video datasets. - Optimize models for latency, throughput, and GPU efficiency using techniques such as quantization, pruning, distillation, batching, and ONNX/TensorRT. Production ML Engineering - Convert prototypes into robust, production-ready services. - Collaborate with backend engineers to deploy models via scalable APIs and micro-services. - Monitor model performance in production and design retraining loops for continuous improvement. Technical Leadership - Establish best practices for experimentation, evaluation, documentation, and reproducibility. - Provide mentorship to junior engineers and contribute to Cliply’s long-term AI roadmap. - Influence architectural decisions across the AI stack to ensure scalability and reliability. Required Qualifications - Bachelor’s degree in Computer Science, Engineering, or related field (Master’s preferred). - 5–10+ years of experience as an ML Engineer, Applied Scientist, or similar role. - Strong proficiency in PyTorch or TensorFlow, with hands-on experience training deep learning models. - Deep expertise in multimodal video machine learning, - Expertise in at least two of the following: - Video understanding / video ML - Computer vision - Speech/audio processing - Natural language processing - Multimodal fusion - Experience with GPU training, distributed training, and large-scale datasets. - Strong understanding of model optimization (quantization, pruning, distillation, ONNX/TensorRT). - Solid software engineering fundamentals (Python, version control, testing, code review). Preferred Qualifications - Experience with multimodal architectures (video-text, audio-text, cross-modal transformers). - Experience with MLOps tooling (MLflow, Weights & Biases). - Prior work in startup environments or fast-paced product teams. - Contributions to open-source ML projects or competitive ML experience (e.g., Kaggle). Why Join Cliply - Build the core intelligence layer of a next-generation video understanding platform. - Own end-to-end architecture and model design - your work becomes the product. - Work with a founder-led team that values technical excellence, autonomy, and speed. - Shape the future of multimodal AI in a real product used by creators and enterprises. Skills:- Video Editing, Multi-modal AI, Computer Vision, Natural Language Processing (NLP), Python, PyTorch and TensorFlow
Get AI-Matched to This Job
Upload your resume and our AI will score how well you match this and thousands of similar roles.