Build and operate scalable batch and streaming data pipelines, connectors, and RAW-to-CUBE transformations using Python, Spark, Databricks, Kafka, dbt, SQL, and Delta Lake. Ensure data quality, schema validation, observability, performance, cost efficiency, and reliable recovery. Develop governed data products and feature-serving interfaces for machine learning models, agents, and downstream applications.
Job Summary
CONVO is seeking a Senior Data Engineer to build reliable data pipelines and connectors powering its CPG-focused agentic platform. You’ll develop scalable batch and streaming pipelines using Python, Spark, Databricks, Kafka, dbt, and Delta Lake, transforming source data into governed RAW and CUBE layers while ensuring data quality, performance, observability, and reliability. You’ll also build data-product and feature-serving interfaces that enable ML models, agents, and downstream applications to consume trusted data efficiently.
Technical mission- Build reliable connectors and RAW-to-CUBE pipelines across batch and streaming paths, and expose governed data and ML features to downstream services.
- Develop Python/Spark ingestion and transformation pipelines from source systems into RAW and Enriched CUBE layers.
- Implement streaming paths alongside scheduled workloads.
- Create dbt/SQL transformations, data-quality controls and contract-validation checks.
- Optimize Delta Lake layout, SQL performance, partitioning and incremental processing.
- Build feature-serving and data-product interfaces for models, agents and applications.
- Instrument pipelines for lineage, freshness, throughput, failure recovery and cost.
- Minimum experience: 5+ years in data engineering, including 3+ years delivering production Spark or python pipelines.
- Advanced Python and production Apache Spark/Databricks development.
- Kafka or comparable event-streaming technology.
- Strong SQL performance tuning and dimensional/data-product implementation.
- dbt and Delta Lake, including incremental patterns and table optimization.
- Batch and streaming reliability patterns: idempotency, checkpointing, replay and late-arriving data.
- Automated data quality, observability and schema-contract testing.
- ML feature stores or online/offline feature consistency.
- CPG, retail, ERP, POS or syndicated-data pipelines.
- Kubernetes-based data workloads and cloud cost optimization.
- Production connectors and RAW-to-CUBE pipelines with automated tests.
- Batch/streaming operational dashboards and recovery procedures.
- Documented data products and feature-serving interfaces.
- Performance and cost baselines for the implemented workloads.
- Works under the Data Architect with source-system owners, ML/optimization teams, platform infrastructure and downstream application teams.
Similar Jobs
Cloud • Information Technology • Security • Software • Cybersecurity
Leads cybersecurity data-source onboarding and integration for Zscaler’s security platform. Builds and automates Python/API-based data transformations, maps and normalizes security data, manages pipeline lifecycle activities, troubleshoots data quality, and identifies security gaps. Partners with cybersecurity SMEs and cross-functional teams to create implementation plans, communicate technical findings, and provide product improvement feedback. The role requires expertise in security platforms, diverse security-tool integrations, data modeling, SQL, and Python.
Top Skills:
Ai/MlAPIsCloud LogsCmdbCnappCspmEdrPythonSIEMSQLUnified Vulnerability Management
Healthtech • Pharmaceutical
Senior Data Engineer responsible for making raw data usable and accessible across the organization. The role involves managing complex data engineering projects, contributing to policies and procedures, recommending improved processes and models, developing data acquisition, integration, and provisioning capabilities, and mentoring less experienced colleagues.
Food • Mobile
Design and build scalable data and AI infrastructure for production LLM applications. Responsibilities include developing RAG systems, embedding and vector retrieval pipelines, AI agents, LLM evaluation and observability frameworks, and batch or streaming data platforms using Databricks, Spark, Snowflake, Delta Lake, and Airflow. The role also involves optimizing distributed workloads, establishing data governance and quality practices, and partnering with engineering and product teams to productionize AI capabilities.
Top Skills:
Ai AgentsApache AirflowSparkCi/CdDatabricksDatabricks Mosaic AiDatabricks Vector SearchDelta LakeEmbeddingsGenerative AiJavaKafkaLangchainLanggraphLlamaindexLlmsMlflowPineconePythonQdrantRagScalaSemantic SearchSnowflakeSpark Structured StreamingSQLUnity CatalogVector DatabasesWeaviate
What you need to know about the Kolkata Tech Scene
When considering the industries shaping India's tech scene, gaming might not immediately come to mind. However, in the last decade, increased internet usage and greater access to mobile devices have catapulted the industry to new heights, with Kolkata-based companies like Virtualinfocom, Red Apple Technologies and Digitoonz, at the forefront, driving the design and animation of new gaming titles for players.



