Senior Manager, Data & Storage Reliability Engineering
Remote · Full Time
Be the first to apply
- Experience
- 10+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 3 days ago
- Work mode
- Work from home
- Education
- Bachelor's degree in Computer Science, Engineering, or related field
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About ServiceNow
The origins of ServiceNow trace back to when engineer Fred Luddy automated a repetitive task for a colleague, inspiring the mission to minimize busywork so people can focus on meaningful efforts. Currently, ServiceNow is recognized as an AI platform that helps 85% of the Fortune 500 companies enhance efficiency and transform business processes by integrating AI, data, and workflows. The company fosters an AI-driven culture blending technology and talent to accelerate innovation.
Role Overview
We are seeking a highly experienced Senior Manager to lead our Data & Storage Reliability Engineering team. This leadership role involves enhancing engineering quality in prevention engineering, reliability, observability, incident analysis, diagnostics, automation, capacity forecasting, and mitigating platform risks. The position requires profound expertise in distributed systems, database architectures, large-scale SaaS environments, and operational excellence, combined with strong leadership capabilities.
Responsibilities
- Design and implement strategic plans for prevention, reliability, observability, resilience, and risk reduction across large-scale production systems.
- Lead transformation of production signals, incident learnings, escalations, migration results, and platform telemetry into effective engineering improvements.
- Collaborate closely with SWAT and Customer & Production Engineering teams to create continuous feedback loops balancing operations and platform enhancements.
- Enhance observability, diagnostics, automation, reliability reviews, resilience validation, migration readiness, and engineering guardrails.
- Identify patterns of failures, reliability risks, observability gaps, inefficiencies, scalability challenges, and performance constraints, then drive corrective actions.
- Set reliability, resilience, observability, automation, and prevention objectives for critical database and storage services.
- Promote proactive monitoring, analytics, and automation to boost operational health and reduce repetitive manual interventions.
- Lead root cause analyses to ensure sustainable solutions for recurring and customer-impacting issues.
- Collaborate with engineering leadership to align database, storage, and platform architecture priorities based on production insights.
- Build and grow a world-class team specializing in reliability, prevention engineering, observability, and platform engineering.
- Manage recruitment, onboarding, career development, goal setting, project prioritization, and performance evaluations.
- Oversee on-call and escalation activities related to production-support operations.
- Cultivate a culture focused on eliminating repetitive manual tasks through automation, self-service diagnostics, guardrails, and scalable engineering.
- Drive cross-team initiatives to enhance the ServiceNow platform’s reliability, scalability, resilience, and operational efficiency.
- Participate in escalation and crisis management, converting immediate recovery actions into durable preventive engineering.
- Evaluate and improve processes to foster continuous enhancement, efficiency, and prevention-oriented engineering methodologies.
- Provide comprehensive training, documentation, dashboards, playbooks, and support to teams interfacing with Data & Storage Reliability Engineering.
- Facilitate integration of new hires, technologies, systems, and automation to scale team effectiveness.
Qualifications
- Experience applying or evaluating AI integration for workflow automation, decision-making, or problem-solving.
- Over 10 years of expertise in database engineering, reliability engineering, distributed systems, platform or infrastructure engineering, production engineering, or managing large-scale SaaS environments.
- At least 4 years of leadership experience managing engineering teams, including cross-functional or distributed groups.
- Proven leadership in teams focused on reliability, database, platform, infrastructure, performance, scalability, or production engineering.
- Strong knowledge of database technologies, OS performance, distributed systems, cloud-native infrastructure, and large-scale production platforms.
- Proficiency in reliability engineering concepts, observability, diagnostics, root cause analysis, capacity planning, scalability, resilience, automation, and operational best practices.
- Experience translating operational telemetry, customer feedback, incidents, and escalations into prioritized engineering actions.
- Skill in designing and enhancing observability systems, diagnostics frameworks, reliability reviews, migration preparations, resiliency validations, automation, and engineering guardrails.
- In-depth tuning and troubleshooting skills across database, OS, storage, network, and application layers.
- Track record of solving complex issues related to reliability, scalability, performance, efficiency, and operational bottlenecks in distributed systems.
- Experience in leading critical investigations of reliability, capacity, performance, scalability, resilience, and operational risks.
- Expertise in leveraging telemetry and observability platforms to analyze and improve system behavior.
- Collaboration experience with software engineering, infrastructure, operations, and escalation units to boost platform reliability and performance.
- Ability to drive engineering efforts using data, metrics, incident analyses, benchmarking, and measurable effects.
- Excellent communication, stakeholder engagement, and leadership capabilities.
- Bachelor’s degree in Computer Science, Engineering, or equivalent technical experience.
Preferred Skills
- Experience operating enterprise-scale database and storage platforms managing critical workloads.
- Background growing teams in Reliability, Database, Platform, Performance, Scalability, or Production Engineering domains.
- Familiarity with observability and telemetry solutions, diagnostics platforms, reliability metrics, incident analytics, and impact reporting.
- Knowledge of reliability reviews, resiliency checks, migration readiness, capacity forecasting, workload simulations, prevention initiatives, and operational risk frameworks.
- Applying AI technologies to enhance anomaly detection, forecasting, incident analysis, prioritization, operational efficiency, and engineering productivity.
- Understanding distributed system architectures, cloud operations, Linux-based production environments, and hyperscale platforms.
- Contributions to platform architecture, database strategy, reliability investments, scalability roadmaps, and long-term engineering development.
- Experience with performance testing, benchmarking, workload simulation, and capacity modeling in support of reliability and prevention engineering programs.
- Working knowledge of enterprise database systems such as MySQL, MariaDB, PostgreSQL, Oracle, SQL Server, or cloud-native databases.
- Awareness of ServiceNow platform architecture and experience with large-scale SaaS operations.
Additional Information
Work Personas
The company supports a flexible and trusting work environment. Employees are assigned work personas based on their job and location, which may include flexible, remote, or office-based categories. Eligibility for work personas may be verified by assessing proximity to the nearest company location.
Equal Opportunity Employer
ServiceNow is committed to equal employment opportunities regardless of race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or legally protected categories. Applicants with arrest or conviction records are also considered per legal guidelines.
Accommodations
The company aims to provide an accessible and inclusive recruitment process. Candidates needing accommodations during the application can request assistance to access alternative application methods.
Export Control Regulations
Employment may be contingent upon gaining necessary export control approvals for access to regulated technology under applicable laws such as the U.S. Export Administration Regulations.
Minimum education
Bachelor's Degree