ServiceNow

Staff Site Reliability Engineer

ServiceNow

Remote · Full Time

Be the first to apply

Experience
6+ yrs
Salary
Openings
1
Posted
1 day ago
Work mode
Work from home
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About ServiceNow

Founded by engineer Fred Luddy, who automated cumbersome work tasks to free employees for meaningful projects, ServiceNow now leads business transformation through AI. Our AI platform integrates any AI, data, and workflow, enabling 85% of Fortune 500® companies to enhance efficiency, speed, and quality. We are cultivating an AI-first culture that integrates innovative technology and talented teams.

Role Overview

As a Staff Site Reliability Engineer, you will design, build, and manage cloud-native engineering platforms that enhance software and release validation, fostering production readiness. You will maintain scalable test and release environments that elevate deployment confidence. Your work will integrate automation, observability, and intelligence into CI/CD workflows, emphasizing productivity and operational excellence.

Key Responsibilities

  • Design and operate cloud-native platforms for software validation and release processes.
  • Maintain production-like test and release ServiceNow environments to improve deployment readiness.
  • Develop automated pipelines incorporating observability, reliability metrics, and quality gates.
  • Create automation solutions that streamline engineering tasks and reduce manual efforts using shift-left engineering principles.
  • Build reusable frameworks, self-service tools, mock services, and manage test data effectively.
  • Enhance Kubernetes platforms for scalable testing infrastructure, release automation, and developer self-service.
  • Implement validation processes for failure detection, policy enforcement, security, resilience, and operational health.
  • Tackle complex challenges in platforms, infrastructure, and networking through software engineering and system design.
  • Collaborate with engineering teams to boost platform reliability, cloud-native adoption, and release quality.
  • Participate in architecture and technical design reviews for scalable automation solutions.
  • Drive technical decisions through quality engineering delivery and strong collaboration.
  • Mentor peers via code reviews, knowledge sharing, and promoting best engineering practices.
  • Champion a culture of reliability, automation, operational excellence, and continuous improvement.

Qualifications

  • Experience integrating AI into workflows, automation, decision-making, or problem-solving.
  • 8+ years in Site Reliability Engineering, DevOps, Platform, Software, or Infrastructure Engineering with a Bachelor's degree; or 6 years with a Master's; or 3 years with a PhD; or equivalent experience.
  • Proficient with Kubernetes cluster operations, networking, storage, security, autoscaling, and multi-cluster management.
  • Skilled in building and managing cloud-native platforms with scalable, high-availability services.
  • Experience integrating Kubernetes with CI/CD pipelines, GitOps, automated tests, deployment validation, and native cloud workflows.
  • Proven automation expertise to enhance developer productivity, release reliability, and operational efficiency.
  • Familiarity with progressive delivery methods like canary releases, feature toggles, automated rollback, and verification.
  • Knowledge of chaos engineering, resilience testing, disaster recovery, and reliability assessments.
  • Strong software development skills in Python, Go, Java, or Ruby, with experience in design, development, and debugging.
  • Understanding of observability tools, monitoring, SLIs/SLOs, incident management, and production operations for distributed systems.
  • Ability to solve complex technical challenges independently and collaborate cross-functionally.
  • A proactive, responsible mindset thriving in evolving, fast-paced environments with enthusiasm for continuous learning and automation.
  • Collaborative, intellectually curious, and effective at working with globally distributed teams.

Preferred Skills

  • Experience with observability and monitoring solutions at scale.
  • Knowledge of DevOps automation and Agile practices utilizing tools like GitLab CI/CD, Argo CD, or Flux.
  • Expertise in enterprise-scale test automation frameworks, e.g., Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit/TestNG.
  • Familiarity with test orchestration, regression testing, flaky test detection, parallel execution, and test data management.
  • Experience with service virtualization, contract testing, synthetic testing, and enabling developer self-service platforms.
  • Use of Infrastructure as Code and configuration management tools such as Ansible or Terraform.
  • Hands-on experience with Kubernetes ecosystem technologies like Helm, Argo Workflows, Kustomize, Istio/Linkerd, Gateway API/Ingress, Prometheus, OpenTelemetry, and container runtimes.
  • Operating Kubernetes platforms in public clouds including AWS (EKS), Azure (AKS), and GCP (GKE).
  • Implementation of progressive delivery pipelines and techniques including automated rollback and verification.
  • Familiarity with AI-assisted engineering for testing, operational automation, or cloud-native environments is advantageous.

Additional Information

Work Personas: ServiceNow classifies employee work models based on their tasks and locations. Eligibility for these categories may be verified through third-party location validation.

Equal Opportunity: ServiceNow provides equal employment chances regardless of race, religion, sex, orientation, age, disability, or any protected status. Candidates with legal records are considered fairly per regulations.

Accommodations: Candidates needing reasonable accommodations during the hiring process can request assistance via the company's global talent team.

Export Controls: Employment for roles involving controlled technologies is contingent on government export approvals.

Level

Mid

Tools & software

Kubernetes required

How they work

Teamwork & Collaboration Problem Solving Adaptability Learning Agility Accountability

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer