Lead Data Engineer
Mastercard · Pune, India
This listing is from the archive and may be closed. Browse the latest experienced jobs for current openings.
Mastercard is hiring a Lead Data Engineer in Pune.
Our Purpose Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Lead Data Engineer
Responsibilities
Data Platform & Orchestration is seeking a Lead Data Engineer to design and build next-generation, cloud-native data platforms supporting Mastercard’s global data ecosystem. In this role, you will lead the development of scalable batch and real-time data pipelines, enabling efficient data processing across Data Lakes and Data Warehouses. You will work at the intersection of data engineering, cloud platforms, and distributed systems, contributing to high-impact initiatives and driving engineering excellence. This role is ideal for someone who thrives in a fast-paced, collaborative environment, enjoys solving complex data challenges, and is passionate about building resilient, high-performance systems at scale.
- Design and build scalable batch and real-time data pipelines using Spark, Kafka, and (preferred) Apache Flink
- Develop robust ETL/ELT frameworks for structured and unstructured data
- Build and optimize data ingestion and transformation pipelines for Data Lakes and Data Warehouses
- Implement stream processing solutions for near real-time use cases
- Ensure data quality, lineage, observability, and governance across pipelines
- Optimize data jobs for performance, scalability, and cost efficiency
- Design and operate cloud-native data platforms on AWS, Azure, or GCP
- Leverage managed services such as S3/ADLS/GCS, EMR/Databricks, BigQuery/Redshift/Snowflake
- Implement Infrastructure as Code (Terraform, CloudFormation, or equivalent)
- Ensure high availability, fault tolerance, and disaster recovery
- Drive cost optimization strategies for large-scale data workloads
- Implement secure data access controls aligned with enterprise standards
Platform & Engineering Excellence
Requirements
- Strong proficiency in Object-Oriented Programming and Design (OOP/OOAD) Java (JDK 8+); Python and/or Go is a plus
- Experience building data services and distributed systems
- Strong understanding of multithreading, scalability, and performance tuning
- Strong hands-on experience with AWS, Azure, or GCP
- Experience with cloud-native data services (S3, ADLS, GCS, Databricks, EMR, BigQuery, Redshift)
- Strong experience with Apache Spark (Core, SQL, Structured Streaming)
- Hands-on experience with Kafka or equivalent messaging platforms
- Experience with real-time processing frameworks (Apache Flink preferred or Spark Streaming)
- Strong understanding of ETL/ELT design patterns and pipeline architectures
- Experience with data formats (Parquet, Avro, ORC)
- Knowledge of data modeling (dimensional modeling, star/snowflake schemas)
- Proficiency in Infrastructure as Code (Terraform, CloudFormation, ARM templates)
- Experience with Docker and Kubernetes
- Solid understanding of cloud networking, IAM, and security best practices
- Experience with workflow orchestration tools (Airflow or equivalent)
- Strong SQL skills and experience with Data Warehouse platforms
About Mastercard
See the company's official careers page for full details, then apply using the button below.