Cox Purtell Staffing Services

Lead Observability Platform Engineer

Cox Purtell Staffing Services

Sydney, New South Wales, Australia · Full Time

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
1 week ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Overview

Join a cutting-edge organization in Sydney operating in high-frequency, low-latency, and mission-critical environments. This role is focused on ensuring speed, accuracy, and reliability across complex systems where downtime is costly. The workplace suits experts familiar with high-stakes fields such as finance, trading, telecommunications, and government defense technology.

Role Summary

As a Lead Observability Platform Engineer, you will be responsible for overseeing an end-to-end observability platform encompassing telemetry collection, ingestion, storage, querying, visualization, alerting, and diagnostics. The platform is critical and should be seamless until an issue arises, requiring your immediate and skilled intervention.

Primary Responsibilities

  • Design, construct, and maintain a shared observability platform covering all aspects of telemetry processing and insights.
  • Develop APIs, services, integrations, and dashboards to simplify adoption and ensure operational reliability of observability tools.
  • Enhance scalability, reliability, performance, and cost-efficiency of high-throughput telemetry systems.
  • Improve developer and operator experiences with self-service tools, best-practice workflows, and practical platform abstractions.
  • Ensure the platform's robustness by anticipating failure modes, applying monitoring solutions, managing incidents, and driving continuous improvement initiatives.

Candidate Requirements

  • Extensive engineering background in site reliability engineering, platform engineering, infrastructure, observability tooling, developer tooling, or distributed systems.
  • Experience with diagnosing failure scenarios, debugging workflows, and maintaining service reliability under pressure.
  • Deep technical knowledge of logs, metrics, traces, events, alerting mechanisms, dashboards, and telemetry pipeline architectures.
  • Proven experience building or managing services, tools, or pipelines that serve multiple engineering teams.
  • Proficiency with modern observability and telemetry tools including metrics, logging, tracing, and time-series databases, and staying current with industry trends.

Preferred Tools and Technologies

Familiarity and hands-on experience with Kafka, Grafana, ELK stack, OpenSearch, ClickHouse, VictoriaMetrics, InfluxDB, Telegraf, Vector, OpenTelemetry, and Prometheus is highly desirable.

How they work

Problem Solving Attention to Detail Initiative Stress Management

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer