Sr. Manager, Data Reliability Engineering (8+ years)
Visa · IN - Bengaluru, India
experiencedIN - Bengaluru, IndiaPosted 29 Aug 2026
Visa is hiring a Sr. Manager, Data Reliability Engineering (8+ years) in IN - Bengaluru.
Responsibilities
- Lead the design, development, deployment, and operation of scalable, reliable, secure, and highly available SRE solutions supporting data platforms, data pipelines, and data infrastructure.
- Oversee the creation and maintenance of automation, monitoring, alerting, observability, incident response, disaster recovery, and reliability engineering practices for data systems across multiple platforms and services.
- Guide teams in troubleshooting and resolving complex production issues involving data pipelines, distributed processing systems, storage platforms, orchestration tools, streaming systems, infrastructure services, performance bottlenecks, and service degradations.
- Drive the adoption of automation, CI/CD, infrastructure as code, quality engineering, observability, reliability testing, and operational excellence practices across the Data area.
- Coach and mentor SRE, infrastructure, platform, or operations engineering teams supporting Data, fostering a culture of continuous improvement, accountability, learning, reliability, data integrity, and technical excellence.
- Collaborate with data engineering, analytics, machine learning, product, security, risk, compliance, QA, infrastructure, and business partners to ensure solutions meet business priorities, customer expectations, regulatory requirements, security standards, data governance expectations, and operational risk controls.
- Ensure adherence to reliability, resilience, security, compliance, access management, disaster recovery, change management, and data protection standards throughout the technology and data lifecycle.
- Lead the implementation of service-level indicators, service-level objectives, error budgets, capacity planning, performance optimization, incident metrics, reliability dashboards, and operational health checks for data platforms and data services.
- Partner with data engineering and platform teams to improve the reliability of batch pipelines, real-time streaming workloads, machine learning pipelines, data ingestion frameworks, transformation jobs, metadata processes, and reporting or analytics services.
- Oversee the development and maintenance of technical documentation, runbooks, architecture diagrams, operational procedures, post-incident reviews, production-readiness checklists, and best practices for code quality, change review, and data platform supportability.
This is a hybrid position. Expectation of days in the office will be confirmed by your Hiring Manager.
Requirements
- 8+ years of relevant work experience and a Bachelor's degree, OR 11+ years of relevant work experience.
- Experience in leading teams in the design, development, and deployment of large-scale engineering solutions.
- Experience with cloud infrastructure, distributed systems, observability platforms, and reliability engineering practices.
- Experience in building and maintaining scalable infrastructure, automation frameworks, incident response processes, and service reliability programs.
- Experience with programming, scripting, and infrastructure tools such as Python, Go, Bash, Terraform, Kubernetes, CI/CD platforms, and version control systems such as Git.
- Experience troubleshooting and resolving issues in complex, high-volume, highly available production environments.
- Experience implementing security, compliance, access control, disaster recovery, and operational risk management standards.
- Experience coaching, mentoring, and developing SRE, infrastructure, platform, or operations engineering teams.
- Experience collaborating with engineering, product, security, infrastructure, and business stakeholders to improve reliability, scalability, and operational excellence.
- Experience developing and maintaining technical documentation, runbooks, post-incident reviews, service-level objectives, operational procedures, and reliability dashboards.
Preferred qualifications
- 9 or more years of relevant work experience with a Bachelor’s Degree, or 7 or more years of relevant experience with an Advanced Degree, or 3 or more years of experience with a PhD.
- Experience supporting large-scale data platforms and data infrastructure in production environments.
- Experience with cloud platforms and services, such as AWS, Azure, or GCP, used to operate reliable and scalable data systems.
- Experience with Site Reliability Engineering practices, including monitoring, alerting, incident management, service-level objectives, error budgets, and post-incident reviews.
- Experience improving the reliability of data pipelines, distributed processing systems, streaming platforms, orchestration tools, storage platforms, and analytics environments.
- Experience troubleshooting complex production issues ac
About Visa
Visa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid. At Visa, you'll have the opportunity to create impact at scale — tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world. Join Visa and do work that matters – t