Site Reliability Engineer
Razorpay Software Private Limited · Bengaluru
Razorpay Software Private Limited is hiring a Site Reliability Engineer in Bengaluru.
Razorpay is one of India’s leading full-stack financial technology companies, powering the way businesses move, manage, and grow money. Founded in 2014 by Harshil Mathur and Shashank Kumar with a simple vision - to simplify payments for Indian businesses - we’ve since grown into a fintech powerhouse driving India’s digital payment revolution. Razorpay powers millions of businesses with a smarter, scalable stack that goes beyond transactions to help them truly build and grow. From building AI-native agentic payments, to AI-assisted fraud detection and real-time risk intelligence to automated re
Responsibilities
- Define SLIs, SLOs, and error budgets for critical payment flows (authorization, capture, refunds, settlements, webhooks) and make them the shared language between product and platform teams.
- Own the release lifecycle for payment services: design progressive rollout pipelines (canary, staged, feature-flagged), automated rollback triggers, and make "can we roll back in under 5 minutes" a launch-blocking question.
- Carry the pager for payment-critical services, lead incident command during outages, and drive blameless postmortems where action items actually ship.
- Eliminate toil through software: build automation for failover, capacity management, load shedding, and degradation so that known failure classes cannot recur.
- Harden payment flows against distributed systems failure modes: retry storms, thundering herds, cascading failures, partial outages of banks and network partners, idempotency violations, and reconciliation gaps.
- Run production readiness reviews for new payment services and hold the line on launch gates using error budget data, not opinion.
- Instrument what matters: design alerting that pages on customer-facing symptoms, not noise, and cut mean time to detection and recovery quarter over quarter.
- Practice failure on purpose: game days, chaos experiments, and failure injection against payment-critical paths.
Requirements
- 10+ years of engineering experience, with at least 5 years operating large-scale distributed systems in production (high QPS, multi-region, or systems where sub-1 percent error rates were business-critical).
- Strong software engineering skills in at least one of Go, Java, or Python. You have built tools and services, not just configured them.
- Deep understanding of distributed systems failure modes and the patterns that contain them: circuit breakers, backpressure, bulkheading, graceful degradation, idempotency.
- Solid fundamentals in Linux internals, networking, and databases under load (replication, failover, connection pool exhaustion, lock contention).
- Hands-on experience designing or significantly improving deployment pipelines: canary analysis, automated rollback, feature flags.
- Genuine on-call ownership: you have carried a pager for systems that mattered, led incidents, and can walk us through a specific outage you handled and what you changed afterward.
- Fluency with modern observability (metrics, tracing, structured logging; e.g. Prometheus, Grafana, OpenTelemetry, Datadog, Coralogix, Clickhouse or similar) and experience reducing alert noise.
- Experience defining SLOs and error budget policies from scratch. Contributions to reliability tooling, open source or internal, that other teams adopted.
- The judgment and communication skills to tell a product team "not yet" with data, and the pragmatism to help them get to "yes" quickly.
Preferred qualifications
- Experience in payments, fintech, banking, trading, or another domain where correctness and money are coupled (transactional consistency, exactly-once semantics, reconciliation).
- Experience with Kubernetes at scale, service mesh, and traffic management.
Razorpay believes in and follows an equal employment opportunity policy that doesn't discriminate on gender, religion, sexual orientation, colour, nationality, age, etc. We welcome interests and applications from all groups and communities across the globe.
About Razorpay Software Private Limited
You will be one of the founding SREs at Razorpay, embedded with the payment platform teams that move money for millions of businesses. Your mandate is to take our payment flows from three nines to four and five nines of availability. In payments, a failed request is not a retry, it is a customer's money in limbo. You will define what reliability means here, build the systems that enforce it, and set the standard every future SRE is measured against.