- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 2 weeks ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Xestro
Xestro is an Australian healthcare software company that supports specialist doctors in managing patients and operating their practices. Our cloud-hosted platform is actively used daily by thousands of healthcare professionals across hundreds of clinics throughout Australia. Based in Maroochydore on the Sunshine Coast near Sunshine Plaza, we recently transitioned to a multi-storey office to fuel our next growth stage.
Role Overview
As a Platform Engineer, you will be part of the Platform team reporting directly to the Head of Engineering with technical guidance from the Lead Site Reliability Engineer (SRE). Your core responsibility will involve managing and enhancing our production infrastructure primarily hosted in AWS, featuring a Linux/Apache/PHP/MySQL environment alongside an evolving Kubernetes-based platform. Your daily tasks will include operating Kubernetes clusters (notably AWS EKS), troubleshooting deployment and infrastructure issues, upgrading systems, monitoring performance, managing capacity, and automating manual processes. Key tools in use include GitLab, ArgoCD, Temporal, and Datadog for deploying, monitoring, and managing background workflow processing. Your role will also extend to contributing towards the comprehensive operation of our AWS ecosystem under the Platform team's remit. Being accountable for live production systems means incident response and thoughtful change management are key, including participation in an on-call rotation.
Candidate Profile
The ideal candidate will bring extensive experience from roles such as platform engineering, SRE, systems administration, or cloud infrastructure operations. Essential qualities include sound judgement developed through hands-on management of production environments and a methodical approach to system reliability and incident resolution.
Core Attributes
- Own the operational health and stability of the systems you manage.
- Plan system changes carefully, considering impacts and recovery plans for failures.
- Systematically troubleshoot issues and maintain composure under pressure.
- Communicate clearly, document insights effectively, and escalate when necessary.
- Adhere to established technical guidelines, review protocols, and change management procedures.
- Continuously propose and implement automation to reduce manual repetitive tasks.
- Commit to constant learning and uphold best security practices in operations.
Required Experiences
- Over 5 years in infrastructure, platform engineering, SRE, or equivalent operational roles.
- Proven expertise managing and resolving issues within production Kubernetes platforms, especially AWS EKS.
- Strong Linux system administration and troubleshooting capabilities.
- Hands-on experience configuring and managing AWS services including VPCs, IAM, EC2, DNS, and networking.
- Proficiency with infrastructure as code tools, particularly Terraform.
- Operational experience with CI/CD pipelines supporting production deployments.
- Skilled in analyzing production incidents leveraging metrics, logging, and monitoring frameworks.
- Competence in scripting languages like Bash or Python to automate tasks.
- Experience in planning and executing system upgrades, maintenance procedures, and managing rollbacks.
Highly Preferred Knowledge
- Familiarity with GitLab CI/CD and ArgoCD tools.
- Proficiency in monitoring and logging solutions such as Datadog including alerting and dashboard management.
- Experience with Temporal or similar workflow orchestration platforms, including handling worker deployments.
- Knowledge of Kubernetes scaling solutions such as EKS Auto Mode and Karpenter.
- Advanced Kubernetes networking, ingress configuration, storage management and workload setup.
- Expertise in high availability architectures, capacity planning, and performance troubleshooting.
- Understanding of AWS security measures and infrastructure access control management.
- Experience supporting containerized workloads alongside traditional Linux and EC2 infrastructures.
Why Join Xestro?
- Stable employment within a well-established and expanding healthcare SaaS company.
- Meaningful work that influences patient care and clinic operations directly.
- A collaborative and progressive engineering culture supported by experienced teammates.
- Empowered ownership of production platforms with opportunities to drive system improvements.
- A focus on continuous learning coupled with strong security operational standards.
- Flexible work options that prioritize onsite collaboration.
- Enjoyment of the Sunshine Coast lifestyle encouraging a balanced work-life integration.