E
EXL
Data Engineer

Data Engineer

On-siteMidData EngineerJust posted
✦Role summaryAI-generated

This Data Engineer will build and maintain ingestion pipelines for six key sources into Fabric's Bronze layer, ensuring append-only Delta tables and full provenance. They will also implement Fabric Mirroring, CDC patterns, and incremental load logic, while enforcing data quality checks and monitoring for the entity resolution engine.

Skills required

About this role

Build and operate the data pipelines that feed the Entity Hub. This role lands all six in-scope sources into Fabric, implements standardization and transformation logic, and maintains the data quality checks and monitoring that the entity resolution engine depends on. Reliable, observable ingestion is the foundation the entire programmed rests on.

Responsibilities

  • Ingestion development — build and maintain pipelines to land the six in-scope sources (Secretary of State, D&B, ARROW, E1, hCue, DocCentral) into the Fabric Bronze/raw layer.
  • Mirroring & CDC — implement Fabric Mirroring for supported structured sources and establish change-data-capture patterns; implement watermark/incremental load logic where mirroring is unavailable.
  • Raw layer management — maintain one Delta table per source on an append-only basis, retaining evidence records and full source provenance.
  • Standardization & transformation — implement name normalization, address parsing and attribute standardization logic in Spark notebooks; support identifier-spine construction.
  • Data quality — implement data quality checks, validation rules, threshold alerts and exception handling; support reconciliation against source.
  • Pipeline operations — schedule, monitor and troubleshoot pipeline runs; investigate failures and performance issues; maintain run documentation.
  • Performance tuning — optimise Spark jobs, Delta file sizes, partitioning and pipeline efficiency to manage Fabric capacity consumption.

Documentation — produce and maintain source-to-target mappings, transformation logic documentation and lineage records

Qualifications

Skill Area Specific Requirements Core Engineering Python, PySpark, advanced SQL, Delta Lake, distributed data processing Microsoft Fabric Data Factory pipelines and Copy Activity, Lakehouse, OneLake, Spark notebooks, Environments, Mirroring, Shortcuts Data Integration Batch and incremental ingestion, CDC patterns, watermarking, reprocessing strategies, schema-on-read for varied formats Data Quality Validation rule implementation, completeness/accuracy checks, alerting, exception workflows, reconciliation Modelling Bronze/Silver/Gold medallion layering, cleansing and conformance, standardization of names, addresses, dates and codes Ops & Governance Pipeline monitoring, lineage and metadata capture, access controls, technical documentation

Must-Have Qualifications

  • 4+ years hands-on data engineering with strong PySpark and SQL
  • Production experience building ingestion pipelines from multiple heterogeneous sources
  • Working knowledge of Delta Lake and medallion/lakehouse architecture
  • Experience implementing incremental loads and CDC-style processing
  • Experience implementing data quality checks and troubleshooting pipeline failures

Nice-to-Have

  • Microsoft Fabric hands-on experience (Mirroring, Copy Jobs, Environments)
  • Exposure to entity/master data standardization (name and address parsing)
  • Familiarity with libraries such as Great Expectations for data quality
  • Experience optimising for Fabric capacity/CU consumption

Key Deliverables Owned

  • Operational ingestion pipelines for all agreed sources
  • Bronze/raw layer with one Delta table per source and CDC retained
  • Standardization and parsing transformation logic
  • Data quality checks, monitoring and exception handling
  • Source-to-target mapping and run documentation

Dual Role / Complementary Skills

Complementary with the Entity Resolution engineering workstream — both are PySpark-on-Fabric disciplines, so this role can cross-train on Splink tuning and candidate-pair generation to provide cover. Also supports the Sr. Data Engineer (Lead) on identifier-spine construction, and can assist the VectorDB Engineer with document/attribute preparation in Phase 2.

Apply on EXL →Opens in new tab
score your resume against this role

Similar open roles

TF
NEW

Senior Staff Software Engineer – Data Platform

Thermo Fisher Scientific·Remote, Indiana, United States of America; Remote, Arizona, United States of America; Remote, Colorado, United States of America; Remote, Georgia, United States of America; Remote, Idaho, United States of America; Remote, Kansas, United States of America; Remote, Michigan, United States of America; Remote, Minnesota, United States of America; Remote, Nebraska, United States of America; Remote, Nevada, United States of America; Remote, New Mexico, United States of America; Remote, North Carolina, United States of America; Remote, Ohio, United States of America; Remote, Oregon, United States of America; Remote, Pennsylvania, United States of America; Remote, South Carolina, United States of America; Remote, Tennessee, United States of America; Remote, Texas, United States of America; Remote, Utah, United States of America; Remote, Wisconsin, United States of America
RemoteStaffData Engineer
$136k – $204k USD
yesterday
JC
NEW

Sr Lead Software Engineer - Data Engineering

JPMorgan Chase·Wilmington, DE, United States
On-siteSeniorData Engineer
yesterday
JC
NEW

Lead Software Engineer- Data Engineer/Pyspark/Databricks

JPMorgan Chase·Houston, TX, United States
On-siteSeniorData Engineer
yesterday
Apply on EXL →
Data Engineer at EXL — Gurugram, India | finddatasciencejobs