I
inDrive
Data Engineer

Data Engineer

On-siteSeniorData Engineerposted recently
Role summaryAI-generated

This role requires designing and maintaining a scalable data pipeline infrastructure on GCP, focusing on real-time CDC, streaming ingestion, and multi-domain data integration for analytics and ML. You’ll architect a layered BigQuery data warehouse while optimizing custom Airflow operators and Kafka-based connectors for high-performance ETL workflows.

Skills required

About this role

We are looking for a Data Engineer to join one of the Data Platform teams that works with the Marketing, Growth, partner, and financial data domains.
You will be working with cutting edge cloud technologies (GCP, AWS, BigQuery, Databricks, K8s) and building a large scale data infrastructure for analytics, machine learning, and streaming/CDC data delivery.

Key Responsibilities

  • Build and operate batch and streaming ingestion into a layered BigQuery DWH (raw → ODS → data marts) using Airflow, Debezium CDC over Kafka with protobuf, Pub/Sub, and Dataflow
  • Integrate external data sources end-to-end — marketing platforms (GA4, AppsFlyer, TikTok/Meta/Google Ads), payment providers, S3 buckets, and third-party APIs — including schema contracts, backfills, and reconciliation
  • Engineer the data platform itself in Python: custom Airflow operators and connectors in a shared ETL framework, Kafka Connect on Strimzi (K8s), Cloud Functions, and API integrations with external providers
  • Build CI/CD and change-management tooling for BigQuery: GitHub-based test-and-approval flows, SQL migration engines (Liquibase/Flyway/Bytebase), sandbox validation, backup and rollback
  • Own reliability and correctness of pipelines: idempotency, deduplication, late-data handling, backfill and replay, freshness monitoring and alerting; write integration and unit tests
  • Drive data governance and compliance: ITGC-compliant change management for BigQuery, IAM and least-privilege access, PII policy tags and DLP, Unity Catalog on Databricks, column-level lineage (OpenMetadata/Dataplex), and disaster-recovery planning
  • Build internal data tools and platform services for agentic workflows with data — Streamlit apps, Slack bots, LLM-based agents and MCP servers that help teams find and use data
  • Support analysts and business teams with data requests, fostering data-driven decision-making across the company
  • Contribute to system design and architecture with the development team

Skills, Knowledge and Expertise

  • Strong practical Python: clean, well-structured, and tested code for services, tooling, and data pipelines
  • Solid software design skills (OOP, modularity, design patterns) — we build platform tools for agentic workflows with data and plan to develop data-related backend services, so well-designed code is highly valued
  • Experience building and operating services in a cloud environment (GCP, AWS or similar): CI/CD, containerization, monitoring and alerting
  • Familiarity with Kubernetes and Terraform — our infrastructure runs on GCP/K8s
  • Hands-on experience with DWH-related tasks (BigQuery or another cloud warehouse) and confident working SQL
  • Clear communication with non-engineering stakeholders — a meaningful share of the work is data requests from analysts and business teams
  • Demonstrated ability to take ownership of technologies or services and proactively contribute ideas to the team

Nice to have

  • Advanced SQL: complex queries, window functions, partitioning, clustering, and cost optimization
  • Experience building reliable pipelines around CDC (e.g., Debezium): idempotency, schema evolution, backfills, and reconciliation
  • Analytical data modeling skills: table grain, facts vs dimensions, slowly changing dimensions, and metric definitions
  • Experience with stream processing frameworks such as Flink or Apache Beam/Dataflow
  • Exposure to data governance and audit compliance (ITGC/SOX), Databricks Unity Catalog, or lineage/catalog tooling (OpenMetadata, Dataplex)
  • Interest in building LLM-based agents and AI tooling for data
Apply on inDriveOpens in new tab

Similar open roles

DT
NEW

Analyst I Data Engineering

Dxc Technology·IND - AP - HYDERABAD
On-siteJuniorData Engineer
today
SK
Sponsored

Land 5x More Interviews - Resume & Strategy

Shaqeeq Khan·Built this board, coached engineers from Netflix, Google, IBM, Amazon. 100+ grads placed.
CV ReviewLive One on One CallsStrategy
Book A Call