- Experience
- 4–7 yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 天前
- Work mode
- In office
- Education
- Degree in Computer Science or related IT field
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Singapore Pools
Singapore Pools, founded on 23 May 1968 by the Singapore government, is a not-for-profit organization providing secure and reliable betting services to counter illegal gambling activities. It contributes significantly to the Tote Board, supporting various social service, community, sports, arts, education, and health initiatives. Since 2004, more than $5 billion has been channeled to these causes, alongside approximately $2 billion in annual taxes and duties paid to the Government. The organization holds the highest certification (Level 4) in responsible gaming awarded by the World Lottery Association since 2012. Staff members actively participate in volunteer efforts benefiting disadvantaged children, youth-at-risk, underprivileged families, the elderly, and environmental conservation.
Job Purpose
The OpsWatch Unit provides enterprise-wide operational monitoring to enhance system resilience, enable proactive interventions, and facilitate timely escalations to management. Supported by AI-driven operations tools and Enterprise Service Management practices, the unit identifies emerging operational issues, assesses risk and business impact, coordinates team mobilization, escalates and notifies stakeholders, supervises remediation efforts, and ensures closure of corrective actions. Additionally, the unit trains the AI-driven Telemetry, Learning and Adaptive System (ATLAS) to improve incident detection, alerting, prevention, and recommendations.
Key Responsibilities
- Develop software and automate processes by creating API-based applications, including integrations with generative AI, to streamline operational reporting and eliminate repetitive manual work.
- Implement Site Reliability Engineering (SRE) best practices addressing observability, availability, performance, and incident response; define, measure, and manage Service Level Objectives and Error Budgets in partnership with product engineering teams; participate in on-call rotations, conduct postmortem analyses, and resolve production bottlenecks.
- Support hybrid cloud infrastructure by guiding product teams toward building resilient applications through strict Infrastructure as Code code reviews.
- Maintain and improve DevOps toolchains by enforcing GitOps workflows and integrating automated security and vulnerability scans into continuous integration and deployment pipelines for applications, containers, and infrastructure deployments.
- Create and manage consolidated operational dashboards by merging telemetry data from multiple monitoring tools and native AWS/Azure metrics.
- Produce actionable financial operations (FinOps) reports to monitor and optimize spending within hybrid cloud environments.
- Plan and execute disaster recovery, backup, redundancy, and capacity strategies while keeping high-quality runbooks and documentation up to date.
Candidate Profile
- Holds a degree in Computer Science, Engineering, Information Science, or related IT field, supported by 4 to 7 years of hands-on experience in software engineering, cloud architecture, DevOps, or Site Reliability Engineering roles.
- Possesses professional certifications such as ITIL, FinOps Certified Practitioner, AWS Certified Solutions Architect (Associate or Professional), AWS Certified DevOps Engineer (Professional), AWS Certified CloudOps Engineer (Associate), Microsoft Certified Azure Administrator Associate (AZ-104), Azure Solutions Architect Expert (AZ-305), or DevOps Engineer Expert (AZ-400).
- Demonstrates deep expertise in building scalable hybrid cloud infrastructures using AWS and Azure, including containerization; proficient in modern hosting and networking design patterns and operational excellence principles.
- Experienced in developing production-quality software primarily in Python, Golang, or Java; skilled in deploying API services and serverless applications on AWS Lambda and Azure Functions.
- Highly skilled in Infrastructure as Code and configuration management tools like Terraform, Ansible, and CloudFormation; experienced with Kubernetes orchestration; familiar with Agile delivery processes and deployment pipeline enforcement through code reviews.
- Proficient in version control systems such as Git, applying advanced branching techniques and GitOps methodologies.
- Strong systems knowledge including designing and managing centralized observability platforms, using generative AI, time-series databases, and monitoring tools like Prometheus, Grafana, Dynatrace, Splunk, and InfluxDB; solid understanding of database administration and network architecture.
- Excellent problem-solving abilities and data-driven decision-making skills; strong communication and interpersonal capabilities for effective stakeholder collaboration.
- Passionate about emerging technology trends impacting cloud engineering and site reliability engineering.
Benefits
- Comprehensive total rewards package
- Health and wellness programs
- Continuous professional development and upskilling opportunities
- Engagement in volunteer and community service initiatives
Level
Entry
Minimum education
Bachelor's Degree