Senior Site Reliability Engineer - Monitoring and Anomaly Detection (Monetization)
GitLab · Bangalore, India
GitLab is hiring a Senior Site Reliability Engineer - Monitoring and Anomaly Detection (Monetization) in Bangalore.
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster. The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, in
Responsibilities
- Design, build, and operate metrics, logs, and traces across the Monetization stack using tools such as Prometheus and Grafana.
- Implement automated detection for billing, data, and event anomalies, and route alerts to the designated feature teams.
- Develop reconciliation and data integrity checks across usage and billing pipelines.
- Define and track service reliability targets (SLOs) and service level indicators (SLIs) to measure reliability, write runbooks, and take part in incident handling.
- Explore artificial intelligence and machine learning techniques to predict system anomalies and accelerate resolution.
- Review merge requests and offer feedback to other Monetization engineers.
- Collaborate with Product, Finance, Support, and other partners to turn operational needs into reliable tooling.
What you’ll bring
- Professional experience with Ruby on Rails.
- A background in site reliability or observability engineering, including monitoring, alerting, SLOs, SLIs, runbooks, incident handling, and tools such as Prometheus, Grafana, and OpenTelemetry.
- Experience building anomaly detection, monitoring, or risk management tooling.
- Exposure to data stores for reporting and insights, especially ClickHouse, and the change data capture and event streaming pipelines that feed them, such as NATS JetStream.
- Experience with billing, financial, or other business-critical systems.
- Experience owning a project from concept to production, including proposal, discussion, execution, and monitoring.
- Clear, concise communication about complex technical, architectural, and organizational problems in English, with the ability to propose thorough, iterative solutions in a remote and largely asynchronous work environment.
- Additional relevant experience includes working knowledge of Python for anomaly detection and data work, and experience with Zuora or Salesforce.
About GitLab
You'll report to the Engineering Manager, Observability, Monitoring, and Integrations. Our team works asynchronously, with regular planning and retrospectives, and we jointly manage deployments, incident management, testing, and safe rollouts. For more on how Monetization works, see [Link: Monetization Sub-department Handbook]. How GitLab Supports Full-Time Employees