Job Tank India
Sprinklr logo

Incident Response Engineer ||

Sprinklr · India - Karnataka - Bangalore

experiencedIndia - Karnataka - BangalorePosted 17 Aug 2026

Sprinklr is hiring an Incident Response Engineer || in India - Karnataka - Bangalore.

Sprinklr is the definitive, AI-native platform for Unified Customer Experience Management (Unified-CXM), empowering brands to deliver extraordinary experiences at scale — across every customer touchpoint.

By combining human instinct with the speed and efficiency of AI, Sprinklr helps brands earn trust and loyalty through personalized, seamless, and efficient customer interactions. Sprinklr’s unified platform provides powerful solutions for every customer-facing team — spanning social media management, marketing, advertising, customer feedback, and omnichannel contact center management — enabling enterprises to unify data, break down silos, and act on real-time insights.

Today, 1,900+ enterprises and 60% of the Fortune 100 rely on Sprinklr to help them deliver consistent, trusted customer experiences worldwide.

Job Description

  • Take end-to-end ownership of production incidents from detection, logging, triage, impact assessment, prioritization, escalation, investigation, mitigation, resolution, recovery, and closure.
  • Act as the primary point of contact and single point of accountability for assigned incidents.
  • Lead P1/P2 and Sev0/Sev1 major incidents, establish incident bridges/war rooms, and ensure structured coordination across Engineering, SRE, DevOps, Infrastructure, Cloud, Database, Network, Product, Support, and other relevant teams.
  • Perform incident triage and determine the appropriate severity based on customer, business, and platform impact.
  • Ensure the right technical owners and subject-matter experts are engaged through the defined escalation matrix.
  • Drive incidents with clear ownership, timelines, action items, escalation paths, and next steps until service restoration.
  • Analyze application failures, service degradation, infrastructure issues, API failures, database issues, distributed-system problems, cloud incidents, and other production issues.
  • Review application logs, monitoring dashboards, alerts, service metrics, and system dependencies to support incident investigation.
  • Develop a good understanding of application architecture, APIs, microservices, databases, distributed systems, cloud infrastructure, and service dependencies to effectively drive technical discussions.
  • Coordinate with technical teams to identify the root cause, contributing factors, mitigation, and permanent corrective actions.
  • Provide timely, accurate, and structured communication to technical teams, business stakeholders, leadership, and customers as required.
  • Communicate incident impact, investigation progress, mitigation status, recovery updates, and next steps clearly to both technical and non-technical stakeholders.
  • Ensure communication follows defined templates, frequency, SLA, and escalation standards.
  • Drive Root Cause Analysis (RCA) and Post-Incident Reviews (PIRs) for major incidents and ensure corrective and preventive actions are assigned, tracked, and closed.
  • Analyze recurring incidents and work with Problem Management and Engineering teams to eliminate repeat issues and improve service reliability.
  • Monitor production systems, service health dashboards, alerts, and operational tools to proactively identify customer-impacting issues.
  • Work with Monitoring, SRE, and Engineering teams to improve alert quality, thresholds, detection coverage, dashboards, and early-warning mechanisms.
  • Identify monitoring gaps and recommend new alerts, dashboards, automation, correlation, and auto-remediation opportunities.
  • Review incident and alert trends to identify opportunities for alert reduction, tuning, proactive detection, and operational automation.
  • Follow ITIL/ITSM-aligned Incident, Major Incident, Problem, and Change Management processes.
  • Coordinate with Change and Release Management teams to identify and minimize production risks associated with releases, maintenance, and infrastructure changes.
  • Support service transition activities and ensure operational readiness before services are moved into production.
  • Develop, maintain, and continuously improve incident management SOPs, runbooks, escalation matrices, response procedures, communication templates, and operational documentation.
  • Track operational metrics including MTTD, MTTA, MTTR, response time, escalation adherence, incident SLA, RCA SLA, customer impact, and repeat incident trends.
  • Prepare incident reports, operational dashboards, management summaries, and trend analysis to drive continuous improvement.
  • Ensure effective shift handovers and continuity of ongoing incidents in a 24×7 global operations model.

About Sprinklr

See the company's official careers page for full details, then apply using the button below.

Apply Now