Build and maintain scalable GCP-based ETL/ELT pipelines using Python, BigQuery, and Google Cloud Storage. Support machine learning infrastructure through feature stores, model data feeds, Vertex AI integration, and MLOps workflows. Ingest data from APIs, streaming platforms, and databases; optimize SQL and BigQuery architecture; implement CI/CD, testing, data quality checks, monitoring, and governance. Collaborate with data scientists, ML engineers, and product managers to operationalize production machine learning systems.
We are looking for someone with advanced Python programming skills who applies robust software engineering principles to data problems. You will collaborate closely with Data Scientists, ML Engineers, and Product Managers to build the scalable, automated pipelines required to train, deploy, and monitor machine learning models in production.
Key Responsibilities
- GCP Pipeline Development: Design, build, and maintain highly scalable ETL/ELT data pipelines using Python and GCP-native data processing tools (e.g., Cloud Run, Cloud Functions).
- AI/ML Infrastructure Support: Engineer feature stores, robust data feeds specifically optimized for machine learning training and inference. Work closely with ML Engineers to operationalize models using Vertex AI.
- Data Integration & Ingestion: Write clean, modular Python code to ingest data from diverse sources (APIs, streaming platforms, on-prem databases) into BigQuery and Google Cloud Storage (GCS).
- System Optimization: Optimize BigQuery architecture, partition/cluster tables, and tune complex SQL queries to ensure performance and cost-efficiency at a massive scale.
- Software Engineering Best Practices: Champion best practices in Python development, including version control (Git), CI/CD pipelines (Cloud Build / GitHub Actions), code reviews, and comprehensive unit/integration testing.
- Data Quality & Governance: Implement robust data quality checks, alerting, and monitoring to ensure the data feeding our AI models is accurate and reliable.
Qualifications
Required Qualifications
- Degree: Bachelor’s or Master’s degree in Computer Science, Engineering, Mathematics, or a related technical field (or equivalent practical experience).
- Experience: 4 to 6 years of professional experience in Data Engineering, Software Engineering, or a closely related field.
- Advanced Python: Deep expertise in Python programming. You should be highly comfortable with:
- Data processing and ML-adjacent libraries (e.g., PySpark, Pandas, NumPy).
- API development
- Writing efficient and production-grade code.
- GCP Mastery: Proven, hands-on experience designing and operating data architectures on Google Cloud Platform. Must have strong experience with:
- BigQuery (advanced SQL, architecture, and optimization).
- Google Cloud Storage (GCS).
- Compute/Serverless (Cloud Functions, Cloud Run).
- AI/ML Acumen: Experience working alongside Data Science teams. A strong understanding of the ML lifecycle, feature engineering, and the data requirements for model training and deployment.
MLOps: Understanding of MLOps principles, model registry, and continuous training pipelines.
Preferred Qualifications
- Vertex AI: Direct experience interacting with or deploying pipelines using Google Cloud's Vertex AI platform.
- Streaming Technologies: Familiarity with real-time data processing using Google Cloud Pub/Sub and streaming Dataflow jobs.
- Infrastructure as Code: Experience managing GCP resources using Terraform.
- Containerization: Proficiency with Docker.
Similar Jobs
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Lead data engineering architecture and execution for large-scale, low-latency analytics. Own technical direction, ensure operational data quality, design streaming and batch pipelines, mentor engineers, coordinate cross-functional teams, and deliver scalable solutions using Spark, Airflow, and AWS or equivalent data platforms.
Top Skills:
AirflowAthenaColumn StoresDatabricksDatabricks ApisEmrFlinkHiveKafkaKappa ArchitectureLambda ArchitectureMaster Data Management (Mdm)Microservices ArchitectureRedshiftSparkSQLStreaming Pipelines
Fintech • Financial Services
Leads data engineering initiatives involving ETL/ELT pipelines, database optimization, orchestration, big data technologies, data modeling, BI performance, and data quality monitoring. Requires strong Scala, Python, SQL, and shell scripting skills, plus payments-domain expertise covering transaction processing, settlements, disputes, reconciliation, gateways, acquiring, issuing, and PCI-DSS compliance. The role also involves cross-functional communication, planning, facilitation, negotiation, and collaboration with senior stakeholders and clients.
Top Skills:
AirflowBusiness IntelligenceData ModelingData Quality FrameworksEltETLHadoopJavaLuigiMonitoringNoSQLPci-DssPrefectPythonScalaShell ScriptingSparkSQLTest Automation
Cloud • Consulting
Design, build, and optimize end-to-end data solutions using Microsoft Fabric, Azure Synapse, Databricks, Azure Data Factory, and Power BI. Develop scalable real-time and batch pipelines, data models, semantic and dimensional models, dashboards, reports, and self-service analytics. Apply medallion architecture, DataOps, monitoring, security, governance, quality controls, and performance tuning. Integrate Azure Machine Learning and streaming technologies when needed while delivering enterprise data and BI solutions for clients.
Top Skills:
Azure Data FactoryAzure Machine LearningAzure Synapse AnalyticsDatabricksDataopsDaxEvent HubsKafkaMicrosoft FabricPower BIPower Query MPythonScalaSpark SqlSQLT-Sql
What you need to know about the Kolkata Tech Scene
When considering the industries shaping India's tech scene, gaming might not immediately come to mind. However, in the last decade, increased internet usage and greater access to mobile devices have catapulted the industry to new heights, with Kolkata-based companies like Virtualinfocom, Red Apple Technologies and Digitoonz, at the forefront, driving the design and animation of new gaming titles for players.


%20(1).png)