Site Reliability Engineer
Sydney, New South Wales, Australia · Full Time
Be the first to apply
- Experience
- 3+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Education
- Bachelor's degree in Computer Science or related field
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Precisely
Precisely stands as a global frontrunner in data quality, enrichment, and location intelligence, empowering highly trusted brands worldwide to make confident decisions based on dependable data. As an AI-first company, artificial intelligence is deeply integrated into our product development, workflows, and problem-solving approaches. Joining Precisely means becoming part of a passionate community of innovators dedicated to leveraging better data for a better world.
Role Overview
The Site Reliability Engineer (SRE) takes charge of the scalability, performance, and availability of Precisely's infrastructure platforms, including the on-premises private cloud managed CEDAR (CCX), the AWS-based RapidCX (RCX) SaaS platform, and the Hosted Managed Services (HMS) AWS cloud deployments. Acting as the bridge between software engineering and operations, the SRE focuses on automation, observability, and reliability standards to maintain operational excellence and platform uptime. The role also involves partnering with engineering teams to embed reliability culture and ensuring production readiness.
Key Responsibilities
- Establish and uphold reliability benchmarks such as SLOs, SLIs, error budgets, alerting systems, logging, tracing, and reduction of mean time to recovery (MTTR) across CCX, RCX, and HMS.
- Develop and sustain infrastructure as code, deployment automation, and observability tools using technologies including Terraform, Ansible, Datadog, Python, and Bash.
- Collaborate with engineering teams to define CI/CD and deployment standards and integrate reliability, scalability, backup, disaster recovery, and fault tolerance principles into service designs.
- Review service designs and conduct Operational Readiness Reviews to ensure readiness for production and disaster recovery scenarios.
- Lead major incident management, including coordinating incident command and communicating effectively with stakeholders.
- Perform root cause analyses, identify repetitive issues, and drive automation initiatives to minimize manual interventions and enhance reliability.
- Maintain detailed runbooks, manage reliability backlogs, and document monitoring gaps, incidents, operational learnings, and risks.
- Leverage Precisely-provided AI tools to aid in infrastructure code development, incident diagnostics, troubleshooting, testing, and documentation.
- Ensure infrastructure compliance with security, vulnerability remediation, data protection, and relevant standards such as SOC 2 and FedRAMP.
- Provide coaching to engineers, partake in cross-team reviews, remain current with SRE best practices, and participate in the on-call rotation for critical escalations and changes.
Minimum Qualifications
- Bachelor's degree in Computer Science, Information Systems, Engineering, or equivalent practical experience accepted.
- At least 3 years of professional experience in systems or infrastructure engineering within enterprise production environments.
- Demonstrated expertise on Linux operating systems (RHEL/Oracle Linux) in multi-site, multi-environment contexts.
- Proficiency using infrastructure-as-code tools such as Terraform and/or Ansible.
- Experienced deploying and managing AWS services, including EC2, ECS, S3, VPC, and IAM.
- Skillful in scripting with Python or Bash for automation solutions.
- Proven track record of building and managing monitoring and alerting frameworks, preferably with Datadog.
- Strong knowledge of TCP/IP networking, DNS, load balancing, and distributed system architectures.
- Experience designing CI/CD pipelines and implementing deployment automation standards.
- Capability to conduct structured root cause analysis and guide post-incident reviews.
- Ability to define service level objectives (SLOs) and facilitate Operational Readiness Reviews to confirm production readiness.
- Effective collaboration skills for working cross-functionally with engineering teams on reliability and observability standards.
- No travel required (approximately 0%).
- Mandatory daily use of Precisely's AI tools such as GitHub Copilot or Claude for coding, troubleshooting, incident analysis, runbook authoring, and documentation.
Preferred Qualifications
- Experience with containerization and orchestration platforms like Docker, ECS, and Kubernetes.
- Familiarity with GitOps processes and source control using Git or GitLab.
- Knowledge of enterprise virtualization in hybrid cloud environments.
- Understanding of change management and ITIL operational methodologies.
- Experience with enterprise security tools such as Qualys, CrowdStrike, or Rapid7.
- Certifications such as AWS Solutions Architect, SysOps Administrator, or DevOps Engineer are an advantage.
Additional Information
Application Integrity Notice: It is unlawful to impersonate another individual during application or interview processes. Any such misconduct can result in application rejection, offer withdrawal, or termination of employment, as permitted by law.
Privacy Statement: Personal data provided during application will be managed in accordance with applicable laws. For further details, candidates should refer to the employer's Candidate Privacy Notice.
Minimum education
Bachelor's Degree