Lead Site Reliability Engineer - Imunify Reliability Platform
Remote · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 1 settimana fa
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Position Overview
This role is with a partner company based in New Zealand seeking a Lead Site Reliability Engineer for the Imunify Reliability Platform. It offers a unique chance to lead site reliability engineering efforts in a greenfield environment focused on a large-scale security product.
You will be responsible for defining what constitutes "healthy" across approximately 70 components including cloud services and agents hosted by customers. Your leadership will facilitate the creation of SLIs, SLOs, error budgets, and standardized monitoring and alerting practices. These efforts will enable early detection of silent degradations in security controls to prevent widespread impact on customers.
Collaboration with engineering leads and senior engineers is critical as you design and develop the telemetry and reliability platform. The environment values measurable results over vanity metrics in a remote-first, asynchronous, and challenging technical setting.
Key Responsibilities
- Define and implement meaningful SLIs for about 70 product components in partnership with squad leads and senior engineers, determining ownership, monitoring, tiering, SLOs, and error budgets.
- Develop a taxonomy for reliability covering service availability, latency, fleet reachability, configuration convergence, security-control effectiveness, artifact delivery, and telemetry pipeline health.
- Guarantee that reliability indicators are independently measurable and resilient to the very failures they monitor.
- Design and build a telemetry collection pipeline suited to customer-hosted agents and cloud services, balancing push-based collection, sampling, privacy, data quality, and data cardinality concerns.
- Collaborate with product teams to extend instrumentation in Python, Go, and Rust components.
- Consolidate current dashboards, queries, and reports into a streamlined observability platform by retiring low-value tooling.
- Implement symptom-based, SLO-driven alerting with multi-window burn-rate methods, defining clear roles for pages, tickets, and dashboards.
- Ensure each production alert has a responsible owner, known failure modes, and actionable runbooks.
- Establish alert maintenance practices including regular reviews, quantitative actionable alert rates, and pruning unnecessary alerts.
- Create a machine-readable ownership and escalation framework to route incidents accurately to engineering squads.
- Define severity levels, expected acknowledgements, follow-the-sun escalation, and clean handoff procedures for multi-time-zone coverage.
- Enhance incident command and blameless postmortems, ensuring reliable timelines, ownership clarity, and effective corrective action follow-up.
- Guide engineering squads to assume ownership of operational duties and on-call responsibilities, discouraging centralized alert buffering.
- Achieve measurable reliability gains within the first year, including full SLI adoption, telemetry deployment, tiered alerts, on-call squad adoption, and quicker detection of security-control degradations.
Qualifications
- Extensive production engineering or site reliability engineering experience with proven ability to establish SLO frameworks beyond operating within existing ones.
- Proficiency in Python with capability to read and modify Go and Rust code for instrumentation and reliability enhancements.
- Hands-on expertise with large-scale time-series and event telemetry systems such as Prometheus/OpenMetrics, Grafana, Alertmanager, and high-cardinality data stores like ClickHouse.
- Experience debugging distributed systems on bare-metal and persistent hosts; not limited to Kubernetes-focused environments.
- Familiarity with production-grade configuration management and CI/CD tools such as Ansible, GitLab CI, Jenkins, or similar.
- Thorough understanding of telemetry challenges related to push-based collection, sampling, clock skew, incomplete reporting, and privacy in customer-managed infrastructures.
- Excellent written and asynchronous communication skills to align teams on measurable system health definitions.
- Strong engineering judgment with a pragmatic approach to observability, alerting, reliability, and operational accountability.
- Experience with security-focused products (e.g., WAF, EDR, antivirus, vulnerability management) is beneficial, especially understanding reliability as effective enforcement rather than just uptime.
- Knowledge of monitoring requirements related to frameworks like SOC 2, ISO 27001, or NIST SP 800-137 is an asset.
- Exposure to OpenTelemetry, eBPF, Sentry, cost-aware telemetry, or cardinality management techniques is advantageous.
- Familiarity with AI-assisted development tools and agentic engineering workflows is a plus.
- Kubernetes experience is useful for smaller platform components running on it.
- Focus of this role is strictly on SRE and reliability engineering, not on DevOps ticket handling, build system ownership, cloud cost management, or acting as on-call for other teams.
Benefits and Work Conditions
- Fully remote work with flexible hours, supporting work from any location worldwide.
- 24 paid vacation days annually plus 10 paid national holidays.
- Unlimited sick leave policy.
- Contribution to private medical insurance.
- Reimbursement for co-working spaces and gym or sports activities.
- Professional growth through challenging projects, learning initiatives, mentoring, and knowledge sharing.
- Possibility to receive rewards for patented innovative ideas.
- Remote-first, asynchronous collaboration spanning multiple time zones.
- Leadership opportunity to define the SRE function, reliability standards, and operational culture from inception.
Level
Lead