I
inDrive
Data Engineer
Data Engineer
On-siteSeniorData Engineerposted recently
Role summaryAI-generated
This role requires designing and maintaining a scalable data pipeline infrastructure on GCP, focusing on real-time CDC, streaming ingestion, and multi-domain data integration for analytics and ML. You’ll architect a layered BigQuery data warehouse while optimizing custom Airflow operators and Kafka-based connectors for high-performance ETL workflows.
Skills required
About this role
We are looking for a Data Engineer to join one of the Data Platform teams that works with the Marketing, Growth, partner, and financial data domains.
You will be working with cutting edge cloud technologies (GCP, AWS, BigQuery, Databricks, K8s) and building a large scale data infrastructure for analytics, machine learning, and streaming/CDC data delivery.
Key Responsibilities
- Build and operate batch and streaming ingestion into a layered BigQuery DWH (raw → ODS → data marts) using Airflow, Debezium CDC over Kafka with protobuf, Pub/Sub, and Dataflow
- Integrate external data sources end-to-end — marketing platforms (GA4, AppsFlyer, TikTok/Meta/Google Ads), payment providers, S3 buckets, and third-party APIs — including schema contracts, backfills, and reconciliation
- Engineer the data platform itself in Python: custom Airflow operators and connectors in a shared ETL framework, Kafka Connect on Strimzi (K8s), Cloud Functions, and API integrations with external providers
- Build CI/CD and change-management tooling for BigQuery: GitHub-based test-and-approval flows, SQL migration engines (Liquibase/Flyway/Bytebase), sandbox validation, backup and rollback
- Own reliability and correctness of pipelines: idempotency, deduplication, late-data handling, backfill and replay, freshness monitoring and alerting; write integration and unit tests
- Drive data governance and compliance: ITGC-compliant change management for BigQuery, IAM and least-privilege access, PII policy tags and DLP, Unity Catalog on Databricks, column-level lineage (OpenMetadata/Dataplex), and disaster-recovery planning
- Build internal data tools and platform services for agentic workflows with data — Streamlit apps, Slack bots, LLM-based agents and MCP servers that help teams find and use data
- Support analysts and business teams with data requests, fostering data-driven decision-making across the company
- Contribute to system design and architecture with the development team
Skills, Knowledge and Expertise
- Strong practical Python: clean, well-structured, and tested code for services, tooling, and data pipelines
- Solid software design skills (OOP, modularity, design patterns) — we build platform tools for agentic workflows with data and plan to develop data-related backend services, so well-designed code is highly valued
- Experience building and operating services in a cloud environment (GCP, AWS or similar): CI/CD, containerization, monitoring and alerting
- Familiarity with Kubernetes and Terraform — our infrastructure runs on GCP/K8s
- Hands-on experience with DWH-related tasks (BigQuery or another cloud warehouse) and confident working SQL
- Clear communication with non-engineering stakeholders — a meaningful share of the work is data requests from analysts and business teams
- Demonstrated ability to take ownership of technologies or services and proactively contribute ideas to the team
Nice to have
- Advanced SQL: complex queries, window functions, partitioning, clustering, and cost optimization
- Experience building reliable pipelines around CDC (e.g., Debezium): idempotency, schema evolution, backfills, and reconciliation
- Analytical data modeling skills: table grain, facts vs dimensions, slowly changing dimensions, and metric definitions
- Experience with stream processing frameworks such as Flink or Apache Beam/Dataflow
- Exposure to data governance and audit compliance (ITGC/SOX), Databricks Unity Catalog, or lineage/catalog tooling (OpenMetadata, Dataplex)
- Interest in building LLM-based agents and AI tooling for data
Apply on inDrive →Opens in new tab
Similar open roles
D
NEWSenior Data Platform Automation Engineer - USA Remote
HybridSeniorData Engineer
$135k – $150k USD
today
SK
SponsoredLand 5x More Interviews - Resume & Strategy
CV ReviewLive One on One CallsStrategy
Book A Call ↗