- Experience
- 5–12 yrs
- Salary
- —
- Openings
- 1
- Posted
- 5 days ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
Epergne Solutions is seeking an experienced SRE DevOps Engineer to join their team in Singapore. This role focuses on supporting and enhancing cloud infrastructure, automation, deployment processes, system reliability, and managing containerized applications. The ideal candidate will collaborate closely with development and operations teams to boost system scalability, availability, and delivery efficiency.
Key Responsibilities
- Design and uphold scalable cloud infrastructure solutions.
- Manage and optimize cloud environments primarily on AWS and GCP; knowledge of Alibaba Cloud is a plus.
- Create and maintain infrastructure using Terraform for Infrastructure as Code.
- Develop and support Kubernetes clusters and containerized workloads.
- Build Python-based automation scripts to increase operational efficiency.
- Implement continuous integration and deployment pipelines along with infrastructure monitoring.
- Investigate and resolve production issues to enhance system reliability and performance.
- Adopt and apply DevOps and Site Reliability Engineering best practices focusing on availability, scalability, security, and incident handling.
- Engage and coordinate with development, infrastructure, and security teams on technical initiatives.
Requirements
- 5 to 12 years of professional experience in DevOps, Site Reliability Engineering, Cloud Engineering, or similar roles.
- Advanced hands-on expertise with AWS and/or GCP cloud platforms.
- Familiarity with Alibaba Cloud is advantageous but not mandatory.
- Strong proficiency with Kubernetes and container technologies.
- Competence in Python scripting for automation tasks.
- Extensive experience with Terraform and managing Infrastructure as Code.
- Proven track record with CI/CD pipelines, monitoring solutions, troubleshooting, and supporting production systems.
- Excellent problem-solving capabilities and analytical thinking.
- Effective communication skills and ability to manage stakeholder relationships.
Skills
Tools & software
Kubernetes
required
Amazon Web Services AWS
required
Google Cloud Platform
· 2 to 5 years required
Terraform
· 2 to 5 years required