Lead Observability Platform Engineer
Sydney, New South Wales, Australia · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Overview
Join a cutting-edge organization in Sydney operating in high-frequency, low-latency, and mission-critical environments. This role is focused on ensuring speed, accuracy, and reliability across complex systems where downtime is costly. The workplace suits experts familiar with high-stakes fields such as finance, trading, telecommunications, and government defense technology.
Role Summary
As a Lead Observability Platform Engineer, you will be responsible for overseeing an end-to-end observability platform encompassing telemetry collection, ingestion, storage, querying, visualization, alerting, and diagnostics. The platform is critical and should be seamless until an issue arises, requiring your immediate and skilled intervention.
Primary Responsibilities
- Design, construct, and maintain a shared observability platform covering all aspects of telemetry processing and insights.
- Develop APIs, services, integrations, and dashboards to simplify adoption and ensure operational reliability of observability tools.
- Enhance scalability, reliability, performance, and cost-efficiency of high-throughput telemetry systems.
- Improve developer and operator experiences with self-service tools, best-practice workflows, and practical platform abstractions.
- Ensure the platform's robustness by anticipating failure modes, applying monitoring solutions, managing incidents, and driving continuous improvement initiatives.
Candidate Requirements
- Extensive engineering background in site reliability engineering, platform engineering, infrastructure, observability tooling, developer tooling, or distributed systems.
- Experience with diagnosing failure scenarios, debugging workflows, and maintaining service reliability under pressure.
- Deep technical knowledge of logs, metrics, traces, events, alerting mechanisms, dashboards, and telemetry pipeline architectures.
- Proven experience building or managing services, tools, or pipelines that serve multiple engineering teams.
- Proficiency with modern observability and telemetry tools including metrics, logging, tracing, and time-series databases, and staying current with industry trends.
Preferred Tools and Technologies
Familiarity and hands-on experience with Kafka, Grafana, ELK stack, OpenSearch, ClickHouse, VictoriaMetrics, InfluxDB, Telegraf, Vector, OpenTelemetry, and Prometheus is highly desirable.