S

Site Reliability Engineer

SGX Group

Singapore · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
19 hours ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

At SGX Group, a trusted global market leader, we build the future of exchange systems, powering markets with robust architectures and platforms. We are seeking a Site Reliability Engineer who approaches operational challenges as software problems, focusing on maintaining production system health through automation, tooling, and workflow orchestration. This role operates within a highly regulated capital markets environment, demanding exceptional reliability, security, and operational discipline.

Responsibilities

  • Maintain and enhance the availability, observability, and recoverability of SGX Group's critical market infrastructure platforms.
  • Establish and uphold standards related to service levels, error budgets, observability, incident response, and automation across engineering teams.
  • Collaborate closely with teams in engineering, infrastructure, security, and product to foster resilience and continuous improvement.
  • Advance Site Reliability Engineering as a core discipline, integrating reliability considerations early in system design.
  • Automate and reduce manual tasks involved in service operation to achieve predictability and stability.

Required Qualifications

  • Over five years of experience in Site Reliability Engineering, platform, or infrastructure engineering with a proven track record of replacing manual processes with automation.
  • Strong programming skills in at least one modern language such as Go, Python, Kotlin, TypeScript, or Rust, producing production-grade code.
  • Hands-on expertise in orchestrating AI-driven operations workflows rather than just utilizing AI for autocomplete.
  • Deep practical experience with Kubernetes, Infrastructure as Code tools like Terraform, CI/CD pipelines, and modern observability tools covering metrics, logs, and tracing.
  • Proven production experience with major cloud platforms, preferably Google Cloud Platform (GCP), with AWS as an alternative.
  • Firm understanding of distributed systems and relevant failure scenarios encountered in live environments.
  • Demonstrated calm and analytical approach during incident response, with thorough root cause analysis and follow-through.
  • Comfort working within complex and regulated operational frameworks.

Preferred Qualifications

  • Knowledge of the FIX protocol or experience in capital markets domain.
  • Experience in building internal developer platforms or self-service tools used by engineering teams.

Why This Role Is Important

You will directly contribute to essential market infrastructure where reliability defines the success of the service. Your decisions will shape how reliability is measured and balanced with business considerations, impacting real-time market operations where downtime is immediately noticeable. This position offers a rare opportunity to build and shape the Site Reliability Engineering discipline from the ground up in a high-stakes environment.

About SGX Group

SGX Group, based in Singapore, is a globally trusted exchange operator known for resilience and transparency. The organization enables price discovery, capital formation, and risk management across various asset classes, supporting a broad ecosystem of issuers, investors, and intermediaries with dependable infrastructure and clearing services.

Tools & software

Kubernetes required Terraform required

How they work

Teamwork & Collaboration Problem Solving Attention to Detail Stress Management

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer