Designs and builds scalable AWS data architectures, ETL pipelines, data lakes, streaming workflows, and API-driven applications. The role uses Lambda, Glue, Spark, Iceberg, Starburst/Trino, Kafka, and AWS orchestration services to integrate financial master and reference data. Responsibilities include schema management, event-driven processing, performance optimization, testing, data format handling, and CI/CD deployment for high-visibility client systems.
We are seeking a highly skilled AWS Data Engineer with deep expertise in AWS cloud architecture, big data processing, real-time streaming, and modern data lake technologies. The ideal candidate will have strong hands-on experience in Spark (PySpark), Iceberg, EMR, Starburst/Trino, and event-driven architectures, along with experience building real-time and API-driven data applications who can design and build generic solutions for one of our Fortune 500 Client programs in the realm of Financial Master & Reference Data Management. This is high visibility, fast-paced key initiative will integrate data across internal and external sources, provide analytical insights, and integrate with the customer’s critical systems.
ResponsibilitiesKey Responsibilities
- Design and implement scalable, secure, and cost-optimized AWS data architectures.
- Develop and maintain ETL pipelines using AWS Lambda and AWS Glue ETL.
- Configure and manage AWS Glue Crawlers, Glue Data Catalog, and schema evolution.
- Build, optimize, and unit test applications on the Apache Spark framework using PySpark.
- Design and optimize data lakes using Apache Iceberg on AWS, including table compaction and Iceberg performance tuning.
- Work extensively with data formats such as Avro, Parquet, JSON, XML, and CSV.
- Orchestrate event-driven workflows using AWS Step Functions and Amazon EventBridge.
- Connect and integrate Starburst from Lambda and Glue ETL jobs for federated querying.
- Implement CI/CD pipelines for automated testing and deployment.
- Perform unit testing using PyTest, and performance tuning of Spark and Python applications
- Strong understanding of AWS architecture best practices, scalability, security, and cost optimization strategies.
- Strong hands-on experience with AWS services including Lambda, Glue ETL, Athena, S3, DynamoDB, Step Functions, EventBridge, SNS, and SQS.
- Deep experience in Apache Spark (PySpark/Scala) development, unit testing, and performance optimization.
- Strong Python programming skills using libraries such as pandas, requests, json, and awswrangler.
- Experience on Apache Kafka and Confluent Kafka.
- Experience designing and optimizing data lakes using Apache Iceberg, including compaction and Iceberg optimization techniques.
Similar Jobs
Information Technology • Database • Consulting
Design and maintain data models, reverse-engineer physical data models, address data integration challenges, and support data architecture strategy. The role requires advanced SQL, ETL, Snowflake, Python, CI/CD, AWS services, and data modeling expertise, with emphasis on AWS S3 and Glue.
Top Skills:
SparkAws AthenaAws EmrAws GlueAws S3Ci/CdETLPythonSnowflakeSQL
Insurance • Software
Build and maintain AWS-based data platforms, scalable pipelines, storage, orchestration, and infrastructure using Terraform. Develop ELT/ETL workflows with Airflow, dbt, and Airbyte; manage Snowflake infrastructure; support containerized services, streaming integrations, observability, data quality, governance, and cost optimization. Use SQL and Python to deliver reliable data products for analytics and machine learning, troubleshoot platform issues, and collaborate with senior engineers and cross-functional teams.
Top Skills:
AirbyteAmazon AuroraAmazon EcsAmazon EksAmazon KinesisAmazon MwaaAmazon RdsAmazon S3Apache AirflowAWSAws IamAws LambdaClaudeCortex CodeDbtDbt CloudDockerGitGithub ActionsKafkaKubernetesPythonSnowflakeSQLTerraform
Information Technology • Consulting • Financial Services
Design, develop, optimize, and maintain scalable ETL pipelines using PySpark and AWS Glue. Build data ingestion and transformation workflows, orchestrate processes with AWS Step Functions, and develop Python-based AWS Lambda functions. Support AWS data lake architectures, analytical use cases, monitoring, reliability, and performance optimization. The role also involves SQL and may include Java microservices, REST APIs, and backend integrations.
Top Skills:
AWSAws Data LakesAws GlueAws LambdaAws Step FunctionsETLJavaPysparkPythonRest ApisSQL
What you need to know about the Kolkata Tech Scene
When considering the industries shaping India's tech scene, gaming might not immediately come to mind. However, in the last decade, increased internet usage and greater access to mobile devices have catapulted the industry to new heights, with Kolkata-based companies like Virtualinfocom, Red Apple Technologies and Digitoonz, at the forefront, driving the design and animation of new gaming titles for players.


