Epergne Solutions

DevOps / Site Reliability Engineer (SRE)

Epergne Solutions

Singapore · Full Time

Be the first to apply

Experience
5–12 yrs
Salary
—
Openings
1
Posted
5 days ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

We are seeking a seasoned DevOps / Site Reliability Engineer (SRE) to join our team in Singapore. The successful candidate will bring 5 to 12 years of experience in managing and automating scalable cloud infrastructure, with a strong focus on AWS, GCP, or Alibaba Cloud platforms. This role requires expertise in Kubernetes, Terraform, and Python scripting for automating and supporting production environments.

Key Responsibilities

  • Architect, deploy, and maintain cloud infrastructure across multiple cloud providers such as AWS, GCP, and Alibaba Cloud.
  • Create and manage Infrastructure as Code (IaC) using Terraform to automate configuration and provisioning.
  • Develop and sustain automation scripts with Python to streamline operational workflows.
  • Operate and troubleshoot Kubernetes clusters and containerized applications to ensure availability and performance.
  • Continuously monitor system health, reliability, and uptime to minimize service interruptions.
  • Establish and maintain CI/CD pipelines to automate software deployments effectively.
  • Work closely with cross-functional development and operations teams to enhance system scalability and robustness.
  • Handle incident response, conduct root cause analyses, and provide production support to resolve issues promptly.

Required Competencies

  • 5 to 12 years of professional experience in DevOps or SRE capacities.
  • Proficient hands-on knowledge of cloud services from AWS and/or GCP; familiarity with Alibaba Cloud is a plus.
  • Advanced skills in Python scripting for automation purposes.
  • Demonstrated expertise managing infrastructure with Terraform.
  • Practical experience with Kubernetes cluster administration and problem resolution.
  • Understanding of CI/CD tooling and best practices in a DevOps environment.
  • Experience monitoring, logging, and optimizing the performance of cloud-native applications.

Preferred Skills

  • Knowledge of containerization technologies like Docker.
  • Insight into cloud networking and security principles.
  • Strong analytical thinking with effective troubleshooting capabilities.

Tools & software

Docker required Kubernetes · 10+ years required Amazon Web Services AWS required GCP · 10+ years required Terraform · 10+ years required

How they work

Communication Teamwork & Collaboration Problem Solving Attention to Detail Work Ethic

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer