Unified Data Platform Architect
This role requires designing and leading a unified data platform at scale for a global ecommerce leader, focusing on seamless data integration and high-performance architectures to support real-time analytics and AI-driven decision-making. The candidate will drive innovation in data infrastructure while ensuring reliability, security, and scalability across diverse markets and user bases.
Skills required
About this role
At eBay, we're more than a global ecommerce leader — we’re changing the way the world shops and sells. Our platform empowers millions of buyers and sellers in more than 190 markets around the world. We’re committed to pushing boundaries and leaving our mark as we reinvent the future of ecommerce for enthusiasts.
Our customers are our compass, authenticity thrives, bold ideas are welcome, and everyone can bring their unique selves to work — every day. We're in this together, sustaining the future of our customers, our company, and our planet.
Join a team of passionate thinkers, innovators, and dreamers — and help us connect people and build communities to create economic opportunity for all.
About the Role
We are looking for a Principal Software Engineer to provide technical leadership for the design and evolution of large-scale data platforms and distributed systems.
This is a senior individual contributor role with responsibility for solving complex, ambiguous engineering problems and establishing technical direction across multiple systems and teams. You will architect platforms that process and serve data at significant scale, combining large-scale batch and streaming computation with reliable backend services and cloud-native infrastructure.
The ideal candidate has deep expertise in distributed systems and data infrastructure, with strong hands-on experience in Java, Hadoop, Apache Spark, Apache Flink, Apache Airflow, and Kubernetes. Python experience is also important for data processing, platform automation, and engineering productivity.
You will remain technically hands-on while influencing architecture, engineering standards, and long-term platform strategy across the organization.
Responsibilities
- Define the architecture and long-term technical direction for large-scale data platforms and distributed processing systems.
- Lead the design of highly scalable, reliable backend and data infrastructure supporting business-critical workloads.
- Architect and build high-performance backend services and platform components primarily in Java, with Python used where appropriate for data processing, orchestration, automation, and tooling.
- Design and evolve large-scale batch and real-time data processing architectures using Apache Spark and Apache Flink.
- Establish architectural patterns for data ingestion, transformation, computation, orchestration, storage, and serving across the data lifecycle.
- Design reliable workflow and dependency-management capabilities using Apache Airflow and related orchestration technologies.
- Define architecture and operational patterns for running large-scale data and backend workloads on Kubernetes.
- Solve complex distributed-systems challenges involving scalability, state management, fault tolerance, consistency, partitioning, backpressure, resource management, and recovery.
- Drive improvements in platform reliability, performance, observability, developer productivity, and infrastructure efficiency.
- Identify systemic bottlenecks and lead architectural initiatives that improve throughput, latency, availability, and cost at scale.
- Establish technical standards and reusable platform capabilities that enable multiple engineering and data teams.
- Lead architecture reviews and provide technical guidance for high-impact initiatives spanning multiple systems and organizational boundaries.
- Partner with architects, senior engineers, engineering leaders, product teams, data engineers, and infrastructure teams to translate business requirements into long-term technical strategy.
- Mentor senior engineers and raise the technical bar through design reviews, code reviews, technical guidance, and engineering best practices.
- Evaluate emerging technologies and make strategic build-versus-buy and architectural decisions for the data platform.
- Lead complex migrations and modernization initiatives while maintaining production reliability and minimizing disruption to dependent systems.
Minimum Qualifications
- 10+ years of software engineering experience, including significant experience designing and operating large-scale distributed systems or data platforms.
- Deep expertise in Java and strong software engineering fundamentals.
- Proficiency with Python for data engineering, automation, or platform development.
- Extensive experience designing and building production backend services and distributed systems.
- Deep hands-on experience with Apache Spark and large-scale distributed data processing.
- Strong experience with the Hadoop ecosystem, including technologies such as HDFS, Hive, and YARN.
- Experience designing and operating real-time or stateful streaming systems using Apache Flink.
- Experience designing large-scale workflow orchestration using Apache Airflow or comparable technologies.
- Strong production experience with Kubernetes, containers, and cloud-native application architectures.
- Deep understanding of distributed-systems concepts including partitioning, replication, consistency, fault tolerance, distributed state, scheduling, resource management, and failure recovery.
- Strong understanding of both batch and streaming architectures and the tradeoffs between different processing models.
- Demonstrated experience driving architecture and technical decisions across multiple teams or major platform initiatives.
- Proven ability to operate effectively in ambiguous problem spaces and turn broad business or platform requirements into executable technical strategies.
- Track record of mentoring senior engineers and influencing engineering practices beyond an immediate team.
Preferred Qualifications
- Experience architecting platforms processing petabyte-scale datasets and/or billions of events per day.
- Deep knowledge of Apache Spark and Apache Flink.
- Experience with modern data lake and lakehouse technologies such as Apache Iceberg
- Experience designing multi-tenant data platforms, including workload isolation, resource governance, capacity management, and cost optimization.
- Strong knowledge of data formats, partitioning strategies, schema evolution, metadata management, and data lifecycle management.
- Experience with data governance, lineage, data quality, and platform observability.
- Experience designing highly available control-plane or platform services supporting large numbers of internal users and workloads.
Core Technology Areas
Languages: Java, Python
Distributed Processing: Apache Spark, Apache Flink
Data Platform: Hadoop, HDFS, Hive, YARN, Iceberg
Streaming: Apache Kafka
Orchestration: Apache Airflow
Infrastructure: Kubernetes, Docker
Architecture: Distributed Systems, Batch & Streaming Processing, Data Lake/Lakehouse, Cloud-Native Data Infrastructu
Additional Details
eBay is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, sex, sexual orientation, gender identity, veteran status, and disability, or other legally protected status. If you have a need that requires accommodation, please contact us at talent@ebay.com. We will make every effort to respond to your request for accommodation as soon as possible. View our accessibility statement to learn more about eBay's commitment to ensuring digital accessibility for people with disabilities.
We use cookies to enhance your experience and may use AI tools for administrative tasks in the hiring process. To learn how we handle your personal data and use AI responsibly, please visit our Talent Privacy Notice, Privacy Center, and AI Hiring Guidelines.