Site Reliability Engineering Technical Lead
Limerick, County Limerick, Ireland · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 17 hours ago
- Work mode
- In office
- Education
- Bachelor's degree in Computer Science, Engineering, or related field
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Company Overview
AMCS Group is a sustainability-focused software specialist headquartered in Ireland with a global presence that includes Europe, the USA, and Australasia. Employing over 1,300 skilled professionals in 22 countries, the company delivers innovative technology solutions aiming to enable a carbon-neutral future.
The company's SaaS products drive efficiency and sustainability improvements across resource-intensive sectors, benefiting more than 5,000 customers in 23 countries by improving profitability and environmental resilience worldwide.
AMCS fosters a growth-oriented and adaptive culture rooted in Irish origins, emphasizing openness, collaboration, creativity, and a strong connection to work, clientele, colleagues, and community.
Role Overview
We are seeking a highly experienced and motivated Technical Lead specializing in Site Reliability Engineering (SRE) and DevOps practices to join our engineering team. This leader will oversee mentoring DevOps engineers, contribute strategically to infrastructure and application architecture decisions, and ensure operational reliability focused on exceptional customer experiences.
Key Responsibilities
- Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs) collaboratively with development and business teams to accurately reflect real customer experiences.
- Lead incident management during complex outages, continuously enhancing detection, diagnosis, and resolution processes by refining alerting systems, tooling, and on-call protocols to minimize mean time to detection (MTTD) and mean time to recovery (MTTR).
- Continuously advance the monitoring and observability infrastructure using tools like Prometheus, Grafana, Mimir, Loki, Tempo, and OpenTelemetry, always prioritizing customer perspective to improve operational effectiveness.
- Conduct blameless root cause analyses and postmortems to translate incidents into long-term improvements, ensuring feedback loops between developers and operations are closed.
- Maintain and enhance high platform availability and performance, proactively identifying and removing bottlenecks before they impact customers.
- Leverage Artificial Intelligence and Large Language Models for automating incident triage, log and trace analysis, runbook execution, and anomaly detection to speed up recovery times and alleviate on-call burden.
- Optimize cloud infrastructure costs by right-sizing workloads, eliminating waste, and designing solutions for cost-efficient scaling within Azure, AWS, GCP, as well as container platforms like Docker and Kubernetes.
- Automate routine operational tasks to reduce toil, enabling platforms to self-heal wherever possible and involve human judgment only in genuine exceptions.
- Provide architectural guidance and take part in design decisions to ensure alignment with organizational goals and best practices.
Success Indicators
- Accurate, actionable, and trusted alerting systems with minimal noise.
- Declining quantity and severity of customer-impacting incidents through effective resolution of root causes.
- Strong bidirectional feedback loops between product engineering and SRE, integrating reliability concerns proactively into product development.
- Significant reduction in repetitive operational work by increasing automation and self-healing capabilities within engineering teams.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, or a related discipline, or equivalent practical experience.
- Over 5 years' experience in DevOps, Site Reliability Engineering, or related roles, including at least 2 years in leadership or mentorship capacities.
- Comprehensive expertise with cloud platforms including Azure, AWS, and GCP, and hands-on cloud architecture experience.
- Demonstrated ability to provide architectural oversight with a focus on performance and scalability.
- Practical experience managing container orchestration environments, especially Kubernetes.
- Proficiency in scripting languages such as PowerShell, Python, or Bash.
- Familiarity with observability tools including Prometheus and Grafana's suite.
- Experience automating systems using tools like Ansible, Terraform, or Chef.
- Strong leadership capabilities, effective communication skills, and a collaborative approach.
Preferred Attributes
- Experience designing and managing CI/CD pipelines using Azure DevOps, Jenkins, GitLab CI, CircleCI, or similar tools.
- Professional certifications such as Certified Kubernetes Administrator (CKA), Certified Kubernetes Application Developer (CKAD), or cloud provider certifications (Azure, AWS, GCP).
- Knowledge of security best practices and regulatory standards for cloud environments.
- Acquaintance with Agile frameworks and project management tools.
Minimum education
Bachelor's Degree