A

Site Reliability Engineering Technical Lead

AMCS Group

Galway, County Galway, Ireland · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
4 days ago
Work mode
In office
Education
Bachelor's degree
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About AMCS Group

AMCS Group is a leading sustainability software provider headquartered in Ireland, operating globally with over 1,300 experts in 22 countries. We deliver innovative SaaS solutions designed to enhance efficiency and promote sustainability across resource-intensive sectors. Serving more than 5,000 clients in 23 countries, our Performance Sustainability software supports profitability and environmental resilience worldwide.

At AMCS, employees experience more than just a job; they have the chance to build meaningful careers in a dynamic and growing company. Rooted in a local ‘start-up’ culture, AMCS promotes openness, collaboration, creativity, and strong connections among people, work, customers, and community.

Job Overview

We are searching for an experienced and motivated Site Reliability Engineering (SRE) Technical Lead to join our engineering team. This role requires a strong grasp of cloud technologies, leadership abilities, and a commitment to operational excellence. The SRE Lead will mentor DevOps engineers, influence architectural decisions, and ensure system reliability focused on delivering outstanding customer experiences. You will collaborate with various teams to maintain scalability, security, and operational robustness of our infrastructure and applications.

Key Responsibilities

  • Define and implement Service Level Indicators (SLIs), Objectives (SLOs), and Agreements (SLAs) in partnership with development and business units, reflecting true customer experience.
  • Lead incident management, improving detection, diagnosis, and resolution speed by refining alerts, tooling, and on-call processes to reduce detection and repair times.
  • Continuously enhance monitoring and observability platforms (e.g., Prometheus, Grafana, Mimir, Loki, Tempo, OpenTelemetry) with a focus on customer needs to boost operational efficiency.
  • Conduct blameless root cause analyses (RCAs) and post-incident reviews to ensure incidents lead to sustainable improvements and tight collaboration between developers and operations.
  • Guarantee high platform availability and responsiveness, proactively identifying and mitigating performance bottlenecks before impacting users.
  • Leverage AI and large language model technologies in operations for incident triage, log/trace analysis, automated runbook execution, and anomaly detection to reduce time to recovery and on-call strain.
  • Optimize cloud resource utilization and cost-effectiveness across platforms like Azure, AWS, and GCP, including container management with Docker and Kubernetes.
  • Develop automation to minimize repetitive tasks (toil), enabling the platform to self-heal and alerting humans only when necessary judgement is needed.
  • Engage in architectural planning and decision-making to ensure infrastructure and systems align with strategic goals and best practices.

Indicators of Success

  • Precise, actionable alerts that team trusts, with minimized noise.
  • Decreasing frequency and severity of production incidents, addressing root causes rather than symptomatic fixes.
  • Continuous two-way feedback between product engineering and SRE to improve reliability and inform product development.
  • Reduced manual operational effort as automation and self-healing capabilities increase.

Required Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
  • Five or more years of experience in DevOps, SRE, or related technology roles, including at least two years in leadership or mentoring positions.
  • In-depth expertise with cloud providers (Azure, AWS, GCP) and practical experience in cloud architecture.
  • Strong background in architectural oversight with proven decision-making skills enhancing system scalability and performance.
  • Experience with container orchestration, specifically Kubernetes.
  • Proficiency in scripting languages such as PowerShell, Python, or Bash.
  • Familiarity with monitoring and logging technologies including Prometheus and the Grafana ecosystem.
  • Hands-on experience with infrastructure automation tools such as Ansible, Terraform, or Chef.
  • Excellent leadership, communication, and teamwork abilities.

Preferred Skills and Certifications

  • Experience designing and managing CI/CD pipelines using tools like Azure DevOps, Jenkins, GitLab CI, or CircleCI.
  • Kubernetes certifications (CKA, CKAD) or cloud certifications (Azure, AWS, GCP) are advantageous.
  • Understanding of cloud security best practices and compliance standards.
  • Working knowledge of Agile practices and project management tools.

Minimum education

Bachelor's Degree

Tools & software

Docker required Kubernetes required Prometheus required Chef required Amazon Web Services AWS required Ansible required Grafana Labs Grafana Cloud required Microsoft Azure required Terraform required

How they work

Communication Teamwork & Collaboration Leadership
🤖
Online · instant AI help
Broxer