Leads the design, development, and maintenance of scalable ETL/ELT pipelines and data products across distributed, multi-cloud environments. Defines data engineering architectures, resolves complex production issues, and designs efficient data solutions. Provides technical leadership and mentorship while collaborating with architects, engineers, and business stakeholders. Requires expertise in big data, AWS services, data warehousing, databases, data modeling, orchestration, streaming, security, observability, version control, containers, and modern development practices.
Job Title Lead Data Engineer Location (s) Cochin Years of Experience 5-10yrs Job Description
Role & Responsibilities Technical Qualifications Python + AWS EMR + AWS Glue + Snowflake + ETL/ELT + Data Engineering + Data Architecture + Big Data + SQL
Lead Data Engineer
What You’ll Do
- Lead the design, development, and maintenance of scalable ETL/ELT pipelines and data products in multi-cloud, multi-region, distributed environments.
- Define and apply appropriate data engineering patterns and architectures based on business and technical requirements.
- Drive technical investigations and resolve complex data, operational, and production issues.
- Design scalable, flexible, efficient, and supportable solutions using appropriate technologies and disciplined development practices.
- Provide technical leadership, guidance, and mentorship to data engineering teams.
- Collaborate with architects, engineering teams, and business stakeholders to deliver robust data solutions.
Key Skills
- Data Engineering & Architecture: Strong understanding of data engineering patterns and modern architectures such as Data Mesh, Data Fabric, Data Lake, and Data Warehouse.
- Big Data & Cloud: Proficiency in Amazon EMR, AWS Glue, Data Lake, and technologies for processing large datasets.
- Programming: Strong proficiency in Python, Java, or Scala, with solid OOP expertise.
- Data Warehousing: Strong knowledge of data warehousing solutions, preferably Snowflake.
- ETL/ELT & Orchestration: Hands-on experience with ETL/ELT pipelines and data orchestration tools.
- Databases: Strong expertise in relational and NoSQL databases.
- Data Modelling: Experience designing efficient OLAP and OLTP data models.
- Streaming: Knowledge of streaming data technologies and architectures.
- Version Control: Experience with Git and modern development practices.
- Containers & Orchestration: Understanding of Docker and Kubernetes.
- Data Security: Knowledge of encryption, access control, data privacy, and compliance.
- Monitoring & Logging: Experience implementing monitoring, logging, alerting, and observability for data pipelines.
Role & Responsibilities
- Lead Data Engineer
- What You’ll Do
- Lead the design, development, and maintenance of scalable ETL/ELT pipelines and data products in multi-cloud, multi-region, distributed environments.
- Define and apply appropriate data engineering patterns and architectures based on business and technical requirements.
- Drive technical investigations and resolve complex data, operational, and production issues.
- Design scalable, flexible, efficient, and supportable solutions using appropriate technologies and disciplined development practices.
- Provide technical leadership, guidance, and mentorship to data engineering teams.
- Collaborate with architects, engineering teams, and business stakeholders to deliver robust data solutions.
- Key Skills
- Data Engineering & Architecture: Strong understanding of data engineering patterns and modern architectures such as Data Mesh, Data Fabric, Data Lake, and Data Warehouse.
- Big Data & Cloud: Proficiency in Amazon EMR, AWS Glue, Data Lake, and technologies for processing large datasets.
- Programming: Strong proficiency in Python, Java, or Scala, with solid OOP expertise.
- Data Warehousing: Strong knowledge of data warehousing solutions, preferably Snowflake.
- ETL/ELT & Orchestration: Hands-on experience with ETL/ELT pipelines and data orchestration tools.
- Databases: Strong expertise in relational and NoSQL databases.
- Data Modelling: Experience designing efficient OLAP and OLTP data models.
- Streaming: Knowledge of streaming data technologies and architectures.
- Version Control: Experience with Git and modern development practices.
- Containers & Orchestration: Understanding of Docker and Kubernetes.
- Data Security: Knowledge of encryption, access control, data privacy, and compliance.
- Monitoring & Logging: Experience implementing monitoring, logging, alerting, and observability for data pipelines.
Similar Jobs
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Design and build data services and pipelines supporting machine learning products. Develop automation tools for model deployment, maintain production data systems, improve infrastructure, participate in code reviews, and collaborate across engineering, data science, and product teams. The role requires expertise in distributed systems, large-scale data processing, CI/CD, container orchestration, and AI-enabled workflow improvements.
Top Skills:
AirflowAWSAws BatchCi/CdDockerEmrGlueGoKafkaKubernetesKv StoresLinuxPythonRelational DatabasesSagemakerSparkSpinnaker
Energy
Develop and maintain data pipelines, models, database structures, and orchestrations on Microsoft Azure. Transform vessel, project, campaign, and third-party data into high-quality datasets and reports. Collaborate with business stakeholders to refine requirements, support data scientists and report developers, and design modern cloud data solutions. The role requires strong data modeling, ER modeling, SQL or programming experience, and at least five years of relevant experience.
Top Skills:
SparkAzure Analysis ServicesAzure Data FactoryAzure DatabricksAzure Synapse AnalyticsData OrchestrationData PipelinesEr ModelingAzureMicrosoft Sql ServerPower BIPythonSQL
Enterprise Web • HR Tech • Professional Services • Software
Implement a large-scale data platform modernization initiative by migrating approximately 1,000 datasets to Snowflake, upgrading Airflow pipelines, migrating Hive to Iceberg, adapting data pipelines, and supporting workflow monitoring. The role requires hands-on Python, SQL, Spark, Kafka, AWS S3, Snowflake, testing, debugging, and large-scale data migration experience. This is a remote contract engagement through March 2027 for candidates in eligible EMEA and APAC locations.
Top Skills:
Apache AirflowApache HiveApache IcebergApache KafkaSparkAws S3ConfluentParquetPysparkPythonSnowflakeSnowpark ConnectSQL
What you need to know about the Kolkata Tech Scene
When considering the industries shaping India's tech scene, gaming might not immediately come to mind. However, in the last decade, increased internet usage and greater access to mobile devices have catapulted the industry to new heights, with Kolkata-based companies like Virtualinfocom, Red Apple Technologies and Digitoonz, at the forefront, driving the design and animation of new gaming titles for players.


