N

Site Reliability Engineer

Nityo Infotech

Singapore · Contract

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
1 hour ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Overview

We are seeking a skilled Site Reliability Engineer (SRE) to join our team in Singapore. This role is ideal for seasoned professionals with over 5 years of experience in Site Reliability, DevOps, or Platform Engineering. The successful candidate will be responsible for building, managing, and enhancing robust and scalable production environments that ensure high availability, reliability, and performance.

Key Responsibilities

  • Guarantee the reliability, availability, and optimal performance of production systems.
  • Design, develop, and manage scalable infrastructure and platform solutions.
  • Create and maintain operational documentation including runbooks, escalation procedures, and support models.
  • Manage incident handling, conduct troubleshooting, identify root causes, and drive continuous operational improvements.
  • Integrate monitoring and alerting solutions with Security Information and Event Management (SIEM) systems and Incident Response processes.
  • Enhance automation efforts, optimize deployment workflows, and boost operational efficiency.

Required Technical Expertise

  • Demonstrable experience in roles such as Site Reliability Engineer, DevOps Engineer, Platform Engineer, Infrastructure Engineer, or Production Engineer.
  • Proficiency with Linux operating systems and shell scripting.
  • Hands-on experience with Git version control and CI/CD pipelines.
  • Expertise in containerization technologies including Docker and Kubernetes.
  • Familiarity with observability platforms and instrumentation including logging, metrics, tracing, alerting, and dashboarding.
  • Experience in integrating enterprise-level monitoring, alerting, SIEM, and incident response tools.
  • Strong knowledge of production environment support and managing incident response.

Candidate Profile

The ideal candidate is passionate about automation, reliability, scalability, and maintaining excellent production operations. Strong analytical and problem-solving skills with a proactive approach to troubleshooting are essential.

Application Information

Interested individuals are encouraged to submit their updated resumes via email or WhatsApp as provided.

Tools & software

Git required Docker required Kubernetes required

How they work

Problem Solving Attention to Detail Work Ethic

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer