Senior Site Reliability Engineer
Dublin, County Dublin, Ireland (Hybrid) · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 56 minutes ago
- Work mode
- Hybrid
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
Merative is undergoing a significant transformation to modernize Micromedex, a vital healthcare solution used worldwide by nearly 3,000 clients. We're establishing an in-house Platform Engineering team and developing Site Reliability Engineering (SRE) expertise, having previously relied on an external operations partner. This senior SRE position will be pivotal in shaping and strengthening our internal SRE and platform engineering capabilities.
This role combines software engineering, application performance, cloud infrastructure, and operations. Unlike a traditional infrastructure or operations position, you will collaborate with engineering and architecture teams to enhance reliability, scalability, and application performance while redefining delivery, observability, and operational standards.
Key Responsibilities
- Establish and evolve SRE practices within the new Platform Engineering team.
- Collaborate with software engineering, architecture, DevOps, and infrastructure teams to boost application reliability and efficiency.
- Serve as a liaison between application developers and infrastructure teams, advocating for non-functional requirements.
- Drive improvements in system performance, scalability, reliability, and operational excellence.
- Design and influence scalable, reliable cloud solutions, mainly focused on Microsoft Azure.
- Enhance and develop modern CI/CD pipelines and deployment processes.
- Strengthen monitoring, logging, alerting, and observability frameworks.
- Define and mature service-level objectives (SLOs) and indicators, advancing beyond general uptime metrics towards detailed reliability insights.
- Lead non-functional and performance testing, as well as reliability improvement initiatives.
- Investigate and resolve complex production and application issues thoroughly.
- Guide engineering teams towards improved development, testing, and operational methodologies.
- Create reusable deployment patterns, automation scripts, and operational runbooks.
- Document and ensure performance and reliability requirements are fully understood before production deployment.
- Coordinate with external operations partners during the transition of expertise to the internal team.
- Contribute to architectural improvements across session management, data and application architecture, APIs, and emerging protocols.
- Take ownership of problem identification, investigation, and resolution end-to-end.
- Support the organization's AI initiatives by exploring safe integration of AI within engineering and operations.
Technology Environment
- Microsoft Azure cloud platform with a multi-region highly available architecture
- Applications built with Java
- Databases: DB2 and Oracle
- Jenkins for automation
- Ansible and custom scripting tools
- Observability tools such as Grafana for monitoring and logging
- Triple-active database architecture supporting operational complexity
The environment continues to evolve with greater adoption of cloud-native, open-source technologies, modern deployment pipelines, and enhanced observability.
Candidate Profile
We seek candidates who combine strong software engineering insights with SRE, platform, cloud, or DevOps experience. Backgrounds may include software engineering, application development, platform engineering, DevOps, or SRE. Special interests should align with application performance, reliability engineering, scalability, cloud architecture, operational troubleshooting, automation, and non-functional testing.
Applicants must be comfortable analyzing application behaviors, not solely infrastructure setups.
Essential Qualifications and Skills
- Proven expertise in Site Reliability Engineering, Platform Engineering, Software Engineering, or related areas.
- Experience collaborating with application development teams and understanding application delivery.
- Deep understanding of application reliability, performance optimization, and scalability.
- Strong troubleshooting skills for complex technical and production challenges.
- Proficiency with cloud environments—Azure experience preferred.
- Designing and implementing scalable, reliable cloud architectures.
- Familiarity with CI/CD pipelines and contemporary delivery methods.
- Good grasp of observability, monitoring, logging, and alerting systems.
- Experience with non-functional, load, or performance testing.
- Knowledge of databases and principles of database scalability.
- Demonstrated problem-solving aptitude with a proactive learning mindset.
- Effective communication skills bridging development and infrastructure teams.
- Influencing technical decisions and driving collaborative, cross-team efforts.
- Hands-on approach to technical problem-solving beyond theoretical or architectural contributions.
Preferred Industry Experience
Experience in healthcare, especially critical and regulated environments like Micromedex, is highly valuable. Other relevant sectors include financial services, enterprise SaaS, and critical software platforms where reliability, security, and operational discipline are paramount.
Success Factors
- Comfort navigating discussions spanning application development and infrastructure engineering.
- Insight into how application design impacts reliability and performance.
- Ability to question existing processes and enhance them effectively.
- Skill in translating reliability requirements into practical engineering solutions.
- Collaboration with software engineers while understanding operational constraints.
- Ownership from problem identification through delivery and embedding of solutions.
- Balancing strategic technical vision with hands-on execution.
Why Join Us?
- Join during a pivotal phase of Micromedex and Merative's engineering transformation.
- Build foundational SRE and platform engineering capabilities internally.
- Influence product architecture, reliability, and engineering workflows.
- Work on a globally-used mission-critical healthcare platform.
- Tackle complex challenges involving performance, scalability, and uptime.
- Engage with a multi-region Azure environment and modern cloud technologies.
- Advance modern delivery and observability techniques.
- Collaborate with a growing engineering team based in Dublin.
- Participate in new platform and product initiatives, including APIs and emerging tech.
- Contribute to the organization's AI strategy and engineering efforts.
- Enjoy a role combining strategic impact and hands-on problem solving.
Additional Information
Candidates must have existing legal authorization to work in Ireland; no visa sponsorship will be provided for this position.
Level
Senior
Industry
Hospitals & Health Care