Data Engineer (Spark/Scala)
Zorba Consulting IndiaHyderabad, Telangana₹1,000,000 – ₹1,700,000
it-jobs
Job Description
We are looking for an experienced Data Engineer (Spark/Scala) with strong hands-on expertise in Apache Spark, Databricks, Scala, PySpark, Python, and SQL . The role involves designing and developing large-scale data pipelines across on-premises and cloud environments , working with multiple file systems and data formats, modernizing legacy workflows, and supporting hybrid data architectures. Key Responsibilities - Design, develop, and maintain scalable data pipelines using Apache Spark, Databricks, Scala Spark, and PySpark . - Build and support complex on-premises data workflows and hybrid on-prem-to-cloud data integration solutions. - Integrate data across HDFS, NAS, on-prem file shares, Amazon S3 , and other storage platforms. - Work with multiple data formats including JSON, Parquet, CSV, Avro, Fixed-Length, and Excel . - Develop optimized SQL queries for data extraction, transformation, and loading. - Connect to multiple relational and non-relational databases and implement performance-efficient data extraction strategies. - Develop and maintain workflow orchestration using Apache Airflow or similar scheduling tools . - Write clean, production-grade Python code for data processing, automation, and engineering utilities. - Develop unit, integration, and data-quality tests for data pipelines. - Troubleshoot pipeline failures, performance bottlenecks, data quality issues, and complex multi-system integration problems. - Support migration and modernization of legacy on-premises data processes to hybrid/cloud environments. - Collaborate with Data Scientists, Analysts, Application Engineers, and other stakeholders. - Create technical documentation covering pipelines, data flows, architecture, and data lineage. - Support cloud integration initiatives, particularly across Azure environments. - Leverage coding assistants and AI agents to improve development productivity and automate engineering tasks. Required Skills - Strong hands-on experience with Apache Spark and Databricks . - Strong experience with Scala/Spark Scala and PySpark . - Strong Python programming skills. - Strong SQL , including complex joins, query optimization, and performance tuning. - Hands-on experience with Amazon S3 . - Experience working with HDFS, NAS, on-prem file systems , and cloud storage. - Strong experience handling JSON, Parquet, CSV, Avro, Fixed-Length, and Excel data formats. - Experience extracting data efficiently from multiple databases. - Strong understanding of complex on-premises data workflows and multi-system integrations. - Experience building hybrid on-prem/cloud data pipelines . - Strong troubleshooting and production support skills. Secondary Skills - Azure cloud services. - Apache Airflow or similar workflow orchestration tools. - Automated unit and integration testing. - Data quality validation and monitoring. - Data lineage and technical documentation. Good to Have - Working knowledge of Java . - Experience with Prefect . - Familiarity with React for internal tools or dashboards. - Experience using AI coding assistants and AI agents . - PBM / Pharmacy Benefit Management / Healthcare domain experience. Preferred Candidate Profile Candidates with strong experience in Scala + Spark + Databricks + PySpark , combined with on-premises data engineering and hybrid cloud integration , will be preferred. Key Skills: Apache Spark, Scala, Spark Scala, Databricks, PySpark, Python, SQL, Amazon S3, HDFS, Airflow, Azure, On-Premise Data Engineering, ETL, Data Pipelines.
Get AI-Matched to This Job
Upload your resume and our AI will score how well you match this and thousands of similar roles.