- Experience
- 4–6 yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Education
- Bachelor's degree in IT or related discipline
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Job Overview
This role is focused on facilitating the effective management and continuous enhancement of Incident Management, Major Incident Management, and Problem Management within designated services. The manager ensures incidents are properly documented, prioritized, escalated, communicated, resolved, validated, and formally closed as per agreed timelines. They investigate recurring or significant problems through root-cause analysis and implement corrective measures to improve service reliability and customer satisfaction.
Key Responsibilities
- Administer and consistently apply Incident, Major Incident, and Problem Management processes, including workflows, templates, escalations, and controls.
- Audit incident records for accurate details including categorization, urgency, priority, timestamps, diagnostic info, resolution, and closure quality.
- Track high-priority, prolonged, breached, reopened, customer-sensitive, and supplier-dependent incidents; proactively follow up and escalate issues as needed.
- Coordinate major incident activities such as declaration, bridge-call facilitation, participant engagement, timeline recording, restoration monitoring, and formal stand-down processes.
- Collaborate across service desk, technical, cloud, infrastructure, application, cybersecurity, supplier, customer, and business units to expedite safe restoration actions.
- Gather and document evidence for major incident reports, post-incident reviews, audits, and customer updates; capture lessons learned and ensure action tracking.
- Identify potential problems from incidents, trends, failed changes, supplier performance issues, and known operational risks.
- Assist in root-cause analyses, document causal factors, and manage corrective/preventive action tracking until closure.
- Manage problem backlogs, known errors, workarounds, resolutions, and maintain related cross-references within ITSM records.
- Ensure timely communication and usage of verified workarounds and known errors to enhance first-line resolution effectiveness.
- Develop operational reports and dashboards monitoring incident trends, SLA compliance, restoration times, backlogs, reopens, major incidents, problem statuses, and action closures.
- Prepare documentation such as meeting packs, minutes, decision logs, and action trackers for reviews and improvement forums.
- Support initiatives to improve monitoring, event management, automation, runbooks, knowledge bases, self-healing, and other workflows aimed at reducing incident detection and resolution times.
- Maintain updated process documents including procedures, priority models, communication templates, escalation lists, and knowledge articles.
- Perform quality reviews, process compliance audits, supplier follow-ups, and assist in closure of identified gaps or corrective actions.
- Identify and recommend improvements to processes, data quality, bottlenecks, supplier delays, and automation opportunities for ongoing enhancement.
Qualifications
- Bachelor’s degree in IT, computer science, software engineering, engineering, information systems, or related fields.
- ITIL 4 Foundation certification mandatory; additional relevant training in incident, major incident, problem management, monitoring, support, or service delivery strongly preferred.
- Additional certifications or training in Site Reliability Engineering, SIAM, COBIT, ISO/IEC 20000, Lean Six Sigma, risk management, business continuity, project management, or information security considered beneficial.
- Comprehensive understanding of IT service management, service operation and transition, service reliability, operational resilience, customer experience, service levels, supplier management, and continuous improvement methodology.
Technical Competencies
- In-depth knowledge of incident, major incident, and problem management lifecycle stages, roles, priority models, escalation processes, communication protocols, governance structures, and recordkeeping standards.
- Proficient skills in auditing incident/problem data for completeness, consistency, SLA adherence, escalation, and closure.
- Experience managing major incident coordination, including communication bridge administration, logging actions and timelines, stakeholder updates, and restoration tracking.
- Capability to support root-cause investigations, differentiate symptoms from causes, document factors, and oversee corrective action progress.
- Familiarity with enterprise ITSM platforms, monitoring/observability tools, knowledge bases, dashboards, workflow automation, configuration management databases, and service mapping.
- Understanding of infrastructure, cloud, networking, telecom, cybersecurity, databases, middleware, applications, and end-user computing sufficient to synergize multi-disciplinary teams.
- Skill in compiling operational dashboards, trend analyses, incident summaries, root-cause reports, meeting records, management presentations, and audit documentation.
- Competence in analytical methodologies such as Pareto charts, 5 Whys, fishbone diagrams, chronology analysis, causal mapping, and foundational data analytics.
Experience
- Between four to six years in IT service operations, particularly incident, major incident, or problem management, or related service desk, reliability, or operational roles.
- A minimum of two years’ hands-on experience coordinating priority incidents, supporting major incidents, maintaining problem documentation, conducting root cause analysis, and monitoring corrective measures.
- Experience utilizing enterprise ITSM tools, monitoring systems, dashboards, knowledge bases, CMDBs, service mapping, and automation workflows.
- Demonstrated ability collaborating across infrastructure, cloud, application, network, cybersecurity, service desk, customer, and supplier teams, often in 24x7 or operationally critical environments.
- Knowledge of operational reporting, customer communications, audit preparation, post-incident analysis, problem backlog management, known-error curation, and continuous improvement efforts is advantageous.
Minimum education
Bachelor's Degree
Skills
How they work
Communication