Site Reliability Engineer
Melbourne, Victoria, Australia · Full Time
Be the first to apply
- Experience
- 3+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 2 weeks ago
- Work mode
- In office
- Education
- Bachelor's degree in Computer Science or equivalent
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Company Overview
Precisely is a leading global provider specializing in data quality, enrichment, and location intelligence. Harnessing artificial intelligence at the core of its innovation, the company collaborates with top global brands to enable dependable data-driven decisions. Joining Precisely means engaging in impactful work at the crossroads of AI and data technology.
Role Summary
The Site Reliability Engineer (SRE) taking this role will be pivotal to the upkeep, performance, and scalability of Precisely's diverse infrastructure platforms, including CEDAR (CCX) for on-premises private cloud services, RapidCX (RCX) as an AWS SaaS cloud environment, and Hosted Managed Services (HMS) also leveraging AWS cloud. This position merges software development and systems operations, focusing on automation, tools for observability, and reliability protocols to ensure platform stability and excellence in operation.
Key Responsibilities
- Define and uphold reliability benchmarks such as SLOs, SLIs, error budgets, alerting, logging, and tracing across platforms like CCX, RCX, and HMS to enhance mean time to recovery (MTTR).
- Create and manage infrastructure-as-code scripts, deployment automation, and monitoring tools using technologies like Terraform, Ansible, Datadog, Python, and Bash.
- Collaborate with engineering teams to standardize CI/CD processes, embedding best practices for reliability, scalability, backup, recovery, and failure planning into service designs.
- Evaluate designs and conduct Operational Readiness Reviews (ORRs) to confirm preparedness for production deployment and disaster recovery.
- Lead responses to major incidents through incident command roles and ensure transparent, timely communication with stakeholders.
- Develop root cause analyses, address recurring problems, and automate preventive measures to reduce manual intervention and bolster reliability.
- Maintain detailed runbooks, reliability backlogs, and collective knowledge bases regarding monitoring shortfalls, incidents, insights, and operational risks.
- Utilize Precisely-provided AI resources to assist with infrastructure coding, incident analytics, troubleshooting, testing, and documentation.
- Ensure compliance with security, regulatory, vulnerability remediation, and data protection standards, including but not limited to SOC 2 and FedRAMP requirements.
- Mentor engineering teams, participate in cross-team reviews, stay up-to-date with best SRE practices, and contribute to the on-call rotation schedule to handle critical escalations and changes.
Candidate Profile
- Education: Bachelor’s degree in Computer Science, Information Systems, Engineering, or equivalent practical experience accepted.
- Experience: Minimum of 3 years in systems or infrastructure engineering within enterprise production environments.
- Technical Skills:
- Advanced proficiency in Linux (RHEL/Oracle Linux) across multiple sites and environments.
- Hands-on experience with infrastructure-as-code tools such as Terraform and/or Ansible.
- Experience deploying and managing AWS services including EC2, ECS, S3, VPC, and IAM.
- Proficiency in at least one scripting language, Python or Bash, to develop automation.
- Proven track record with monitoring and alerting implementations, preferably using Datadog.
- Solid grounding in TCP/IP networking, DNS, load balancing, and distributed systems.
- Experience designing CI/CD pipelines and deployment automation standards.
- Strong analytical capabilities for performing root cause analysis and post-mortem reviews.
- Ability to establish SLOs and lead Operational Readiness Reviews, collaborating effectively with engineering teams.
- Track record of cross-functional collaborations on reliability and observability.
- AI Tool Usage: Daily active use of AI coding assistants such as GitHub Copilot or Claude for coding, troubleshooting, and documentation is a requirement, supporting development of infrastructure-as-code, incident investigation, and runbook creation.
- Preferred but not mandatory skills include expertise in containerization and orchestration (Docker, ECS, Kubernetes), GitOps workflows, virtualization in hybrid cloud contexts, ITIL and change management practices, security tools like Qualys and CrowdStrike, and relevant AWS certifications.
- Travel: None required.
Additional Information
Impersonation of another individual in the job application or interview process is strictly prohibited and may result in application rejection, offer rescission, employment termination, and possible legal consequences. Personal data handling complies with applicable laws in accordance with the Precisely Candidate Privacy Notice.
Minimum education
Bachelor's Degree