- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 3 hours ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
We are representing a partner company seeking a highly experienced Staff Site Reliability Engineer based in Ireland for a fully remote position. This senior engineering role involves leading reliability efforts for complex production environments powered by AI technologies. As a technical authority, you will design scalable and resilient infrastructure and implement consistent SRE methodologies to support rapid organizational growth.
The role will encompass responsibilities spanning cloud platforms, Kubernetes orchestration, observability solutions, CI/CD pipelines, data systems, and machine learning operations. You will collaborate extensively with platform, product, data, and ML engineering teams to enhance system availability, security, performance, and efficiency.
Key Responsibilities
- Architect, deploy, and maintain scalable, secure production environments, primarily leveraging AWS infrastructure.
- Lead initiatives to improve reliability practices across diverse engineering teams.
- Design, migrate, optimize, and harden Kubernetes-based infrastructure for production readiness and scalability.
- Implement and enforce infrastructure as code standards, particularly using Terraform or similar technologies.
- Define and operationalize service-level indicators (SLIs), objectives (SLOs), error budgets, and other reliability metrics.
- Enhance observability across applications, infrastructure, data pipelines, and ML models to provide comprehensive insight into system health.
- Collaborate with product and data teams to integrate telemetry and analytics into reliability assessments.
- Design and refine CI/CD pipelines throughout all development stages to increase deployment safety, frequency, and operational predictability.
- Lead incident management for complex, multi-system failures and facilitate rigorous post-incident reviews and improvements.
- Reduce manual operational workload through automation and improved tooling.
- Create processes and infrastructure to standardize, monitor, and troubleshoot customer environments at scale.
- Support production readiness of ML workloads by applying MLOps best practices for deployment, monitoring, and retraining.
- Ensure compliance with enterprise security, regulatory, and operational standards.
- Mentor engineering staff to elevate organizational reliability and engineering capabilities.
- Partner with senior architects and engineers to shape overall technology strategy and product architecture.
Required Qualifications and Experience
- Extensive hands-on experience in Site Reliability Engineering or equivalent infrastructure roles.
- Proven track record establishing or scaling SRE capabilities in complex, high-growth, or distributed environments.
- Advanced proficiency with AWS or Azure cloud platforms and modern cloud-native architectures.
- Deep Kubernetes expertise including migrations, scaling, optimization, and security enhancements.
- Strong skills in Infrastructure as Code, especially Terraform or comparable tools.
- Experience designing and managing CI/CD pipelines across the software lifecycle.
- Comprehensive knowledge of observability tools and practices for distributed systems.
- Ability to troubleshoot and maintain multi-tenant or enterprise-scale customer environments.
- Familiarity with production data platforms and machine learning systems support.
- Direct experience with MLOps workflows including model deployment and monitoring.
- An understanding of distributed systems engineering principles such as scalability, fault tolerance, and resilience.
- Excellent communication and cross-functional collaboration abilities.
- Experience with global-scale B2B or B2C product infrastructures preferred.
- Knowledge of AI/ML, NLP, or large language model-based systems is a significant plus.
- Experience integrating analytics and operational metrics into reliability monitoring.
- Background working in regulated enterprises with strong security and compliance requirements favored.
- Proven ability to scale infrastructure during rapid growth and evaluate technology vendors and platforms.
- Experience supporting large enterprise customers, including complex virtual private cloud deployments, is beneficial.
- Demonstrated problem-solving skills, high ownership, accountability, and proactive risk management.
- Mindset oriented towards continuous learning and persistent process and system improvements.
Benefits and Work Environment
- Permanent, full-time employment.
- Fully remote working arrangement within European time zones.
- Engagement with advanced AI and agent-driven infrastructure technologies.
- Opportunities for significant influence over reliability standards and engineering direction.
- Broad exposure to cloud infrastructure, Kubernetes, distributed systems, data engineering, and machine learning operations.
- Collaboration with highly skilled engineers, architects, product managers, and AI/ML experts.
- Chance to impact technology strategies at scale within a dynamic and growing organization.
- A culture that emphasizes ownership, excellence, and continuous advancement.
- International distributed team environment.
- Potential for career growth and elevated technical leadership roles.
Level
Mid