Job Description
This role is for one of the Weekday's clients Min Experience: 3+ years Location: Bengaluru JobType: full-time We are looking for a highly skilled Big Data Engineer with 3+ years of experience in building scalable data pipelines and distributed systems. The ideal candidate will have strong expertise in Apache Spark (Scala), experience working across on-premise and AWS environments, and a solid understanding of large-scale data processing in AdTech ecosystems. This role involves working on high-volume datasets (billions of records), optimizing distributed jobs, and contributing to the design of robust data infrastructure powering analytics and identity-driven use cases. Requirements Key Responsibilities: - Design, develop, and optimize large-scale batch data pipelines using Apache Spark (Scala) - Process and transform high-volume datasets (TBs of data) in distributed environments - Build and maintain data pipelines across hybrid infrastructure (On-Prem + AWS) - Work with object storage systems such as S3 and MinIO for efficient data access and storage - Develop reusable and scalable data processing frameworks - Optimize Spark jobs for performance (memory tuning, partitioning, shuffling, etc.) - Manage and orchestrate workloads using HashiCorp Nomad - Integrate data pipelines with PostgreSQL and other downstream systems - Ensure data quality, consistency, and reliability across pipelines - Troubleshoot production issues and perform root cause analysis - Contribute to system design discussions, especially for high-scale AdTech use cases (identity resolution, user profiling, etc.) Required Skills: - Strong programming experience in Scala - Good working knowledge of Python (for auxiliary tasks, scripting, or ML integration) - Deep expertise in Apache Spark (Core and SQL) - Strong understanding of distributed data processing and large-scale systems - Experience working with AWS (S3, EMR or equivalent ecosystem) and on-prem clusters - Hands-on experience with object storage systems (S3 / MinIO) - Experience with HashiCorp Nomad or similar orchestration tools - Solid understanding of data modeling and ETL pipeline design - Experience working with PostgreSQL or similar relational databases - Strong debugging and performance tuning skills for Spark jobs - Familiarity with Unix/Linux environments and shell scripting Good to Have: - Experience in AdTech, Identity Graph, or User Profiling systems - Exposure to machine learning pipelines or feature engineering workflows - Experience with data lake architectures - Understanding of cost optimization and resource management in AWS Tech Stack Summary: - Languages: Scala (Primary), Python (Secondary) - Processing: Apache Spark (Core and SQL) - Infrastructure: AWS and On-Prem - Storage: S3, MinIO - Orchestration: HashiCorp Nomad - Database: PostgreSQL Must-have skills Spark, SQL, Scala Good-to-have skills Python, AWS, Big Data
Get AI-Matched to This Job
Upload your resume and our AI will score how well you match this and thousands of similar roles.