N

Senior Product Manager, Observability

Nscale

London, England, United Kingdom · Full Time

Be the first to apply

Experience
5–8 yrs
Salary
—
Openings
1
Posted
2 days ago
Work mode
In office
Education
Degree in Computer Science, Engineering, or equivalent experience
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Overview

Nscale is innovating the AI cloud platform space by building a fully integrated GenAI cloud platform encompassing data centres, software, and applications using sustainable technologies. The company values a culture rooted in innovation, personal ownership, transparency, and collaboration where agility and resilience are prioritized.

Role Summary

As a Senior Technical Product Manager focused on Observability, you will own and advance the part of the Nscale platform that provides real-time insight into GPU fleet operations. This includes managing the telemetry collection pipeline, data aggregation and storage, and the front-end observability interfaces such as logs, metrics, and traces. You'll collaborate closely with teams across software, network, operations, and customer support to ensure fleet health is visible, actionable, and reliable as Nscale expands globally.

Key Responsibilities

  • Lead the strategy and roadmap for the observability platform encompassing telemetry pipelines, data aggregation, trace collection, and customer APIs and dashboards.
  • Specify how telemetry data (logs, metrics, traces) is sourced from physical infrastructure, aggregated, and made accessible for fleet management and incident response.
  • Develop and optimize alerting systems to focus on critical signals, reduce noise, and ensure timely notification to the right stakeholders.
  • Continuously gather and prioritize new telemetry requirements to expand monitoring coverage as the deployment grows across new hardware and sites.
  • Participate in incident reviews and operational workflows to identify manual effort inefficiencies and incorporate enhancements into the platform.
  • Define, track, and improve vital metrics such as alert accuracy, detection and resolution times, telemetry coverage, and platform uptime.
  • Mentor junior product managers, elevating product documentation standards and decision-making quality within the team.

Required Qualifications and Experience

  • 5 to 8 years of product management experience with a demonstrated focus on observability, infrastructure, or operations-centric products.
  • Proven ownership of observability solutions capable of capturing and displaying logs, metrics, and traces at scale, including understanding architectural and UX trade-offs.
  • Practical hands-on familiarity with Prometheus, Loki, Mimir, Datadog, Grafana, or OpenTelemetry.
  • Experience with deployment tooling related to data centre infrastructure, such as provisioning workflows, networking automation, or zero-touch deployments.
  • Background in creating operator-focused solutions for roles including design engineers, project managers, site reliability engineers, and technicians, with a strong enthusiasm for their workflows.
  • Robust technical knowledge allowing leadership in architecture and trade-off discussions involving telemetry pipelines, time-series data stores, alerting mechanisms, and integrations.
  • Track record of converting ambiguous operational challenges into delivered, measurable improvements in visibility, incident handling, or reliability.
  • Exceptional communication skills for engaging engineers, operators, and stakeholders at all levels.

Preferred Skills

  • Experience across broader observability toolsets beyond the primary ones listed.
  • Knowledge of bare-metal provisioning (e.g. OpenStack Ironic, MAAS) or network automation tools (e.g. NetBox, Nautobot).
  • Educational background in computer science or engineering, or previous roles as engineer, SRE, or infrastructure operator.
  • Familiarity with GPU/accelerated computing infrastructure, hyperscale operations, or data centre deployments at scale.
  • Experience with IT Service Management platforms like Jira Service Management, ServiceNow, Zendesk, or Freshservice.
  • History of working in fast-paced, high-growth environments where products evolve alongside the expanding infrastructure.

Equal Opportunity and Inclusivity

Nscale is dedicated to maintaining an inclusive, equitable, and diverse workplace. Applicants of all backgrounds, including people of color, those from the LGBTQ+ community, individuals with disabilities, neurodivergent persons, parents, carers, and socio-economically disadvantaged groups are highly encouraged to apply. The company welcomes requests for accommodations to support applicant needs.

Additional Information

The outlined responsibilities provide a general overview and the role may include other tasks consistent with the role’s qualifications as assigned by management. Candidate privacy is maintained according to company policies.

Minimum education

Bachelor's Degree

Tools & software

Prometheus required Grafana Labs Grafana Cloud required

How they work

Teamwork & Collaboration Adaptability Leadership Accountability Integrity
🤖
Online · instant AI help
Broxer