Staff Software Engineer - AI Reliability Engineering
Dublin, County Dublin, Ireland · Full Time
Be the first to apply
- Experience
- Any
- Salary
- EUR 235,000 – EUR 295,000 / year
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Education
- Bachelor's degree or equivalent experience
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Anthropic
Anthropic is dedicated to developing AI systems that are reliable, interpretable, and controllable. Our objective is to create AI technology that is safe and beneficial for individuals and society at large. Our rapidly expanding team includes researchers, engineers, policy professionals, and industry leaders collaborating to build beneficial AI systems.
Role Overview
You will be instrumental in ensuring the trustworthiness and robustness of Claude, our flagship AI assistant. The AI Reliability Engineering (AIRE) team works collaboratively across various departments at Anthropic to bolster the dependability of critical serving components—from the software development kit (SDK) through network infrastructure, API layers, service architecture, and hardware accelerators. This role involves active engagement with partner teams during incidents and project cycles to fortify system resilience.
Key Responsibilities
- Establish and maintain appropriate Service Level Objectives (SLOs) for large language model serving platforms, optimizing the balance between uptime, responsiveness, and development efficiency.
- Create and deploy comprehensive monitoring and observability solutions along the entire token processing path.
- Design and build fault-tolerant, multi-region serving infrastructure spanning several cloud providers.
- Take lead during critical incidents involving AI services, facilitating rapid recovery, conducting thorough post-incident analyses, and implementing systematic improvements.
- Ensure the reliability of safeguard model deployments, a vital aspect for both site stability and Anthropic’s safety standards.
Qualifications and Ideal Candidate Attributes
- Experience in distributed systems, infrastructure engineering, or site reliability engineering with a focus on robustness and scalability.
- Adaptability and willingness to engage with unfamiliar technical environments during critical situations to help resolve issues efficiently.
- Systemic thinking that recognizes how different components integrate and interact within complex architectures.
- Ability to foster strong cross-team collaboration by building trusted partnerships rather than acting as external consultants.
- A user-focused mindset with a sense of ownership over system outcomes regardless of direct ownership.
- Excellent communication and teamwork skills, preparing you to work across all company functions.
- Diverse technical background, potentially including experience in product stack development, database scaling, or managing large distributed systems.
Preferred Experience
- Roles involving site reliability or production engineering on large-scale infrastructures.
- Handling extensive infrastructure for model serving or training environments with over 1000 GPUs.
- Familiarity with ML-specific hardware accelerators like GPUs, TPUs, or Trainium.
- Understanding of advanced ML network techniques such as RDMA and InfiniBand.
- Expertise in AI-related observability frameworks.
- Experience with chaos engineering and resilience testing methodologies.
- Contributions to open-source ML tooling or infrastructure projects.
Compensation and Logistics
The salary range for this role is approximately €235,000 to €295,000 per year.
Minimum educational requirement is a bachelor's degree or an equivalent blend of education and experience, with relevant technical coursework or professional background.
Our office attendance policy requires staff to be physically present at least 25% of the time, with some roles possibly requiring more on-site involvement. Visa sponsorship is available with dedicated legal support.
Diversity and Inclusion
We encourage applications from all qualified individuals, including those from underrepresented groups. We value diverse perspectives for the ethical and societal impact of the AI technologies we build.
Security Reminder
Only official Anthropic recruiters contact candidates from @anthropic.com emails or authorized agencies. We never request money or financial information before employment begins.
Why Join Anthropic?
We focus our efforts on large-scale, multidisciplinary AI research with the goal of creating steerable and trustworthy AI. We operate as a unified team tackling significant scientific challenges, with a strong emphasis on collaboration and communication.
Our headquarters are in San Francisco, offering competitive pay, benefits, flexible schedules, generous leave policies, and equity donation matching.
Level
Mid
Minimum education
Bachelor's Degree