Staff Software Engineer, AI Reliability Engineering
Dublin, County Dublin, Ireland · Full Time
Be the first to apply
- Experience
- Any
- Salary
- EUR 235,000 – EUR 295,000 / year
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Education
- Bachelor’s degree or equivalent experience
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Anthropic
Anthropic is dedicated to developing AI systems that are dependable, understandable, and controllable with the aim of ensuring AI benefits users and society as a whole. Our expanding team consists of researchers, engineers, policy specialists, and business professionals working collaboratively to build responsible AI technologies.
Role Overview
We seek an engineer to enhance reliability for Claude, our AI assistant, by engaging cross-functionally to improve the robustness of critical infrastructure spanning SDKs, network, APIs, and hardware accelerators. The role demands a holistic perspective on complex distributed systems impacting AI service delivery.
Key Responsibilities
- Establish Service Level Objectives for large-scale language model serving, balancing availability, latency, and development speed.
- Develop and deploy comprehensive monitoring and observability solutions along the AI token processing pipeline.
- Contribute to the design and rollout of fault-tolerant serving infrastructure across multiple regions and cloud environments.
- Lead response efforts during critical AI service incidents, promoting rapid recovery, detailed incident analyses, and continuous reliability improvements.
- Support safeguard model serving reliability, crucial for site stability and safety commitments.
Desired Background and Qualifications
- Extensive experience in distributed systems, infrastructure, or reliability engineering roles.
- Proactive and courageous approach to resolving unfamiliar system incidents.
- Ability to comprehend system integration holistically and identify potential vulnerabilities.
- Strong interpersonal skills for effective collaboration and team integration.
- User-focused mindset with ownership over system outcomes beyond direct responsibilities.
- Excellent communication aptitude for cross-company partnerships.
- Diverse technical experiences ranging from product development to scaling distributed systems.
Additional Preferred Expertise
- Prior roles as SRE, Production Engineer, or equivalent in high-scale system reliability.
- Hands-on experience managing large-scale AI model serving or training platforms (>1000 GPUs).
- Knowledge of ML hardware accelerators such as GPUs, TPUs, and Trainium.
- Familiarity with ML networking techniques like RDMA and InfiniBand.
- Expertise in AI-specific monitoring and observability frameworks.
- Background in chaos engineering and systematic testing to enhance resilience.
- Contributions to open source projects related to infrastructure or ML tooling.
Compensation and Logistics
The annual salary for this position ranges from €235,000 to €295,000.
Minimum education requirement: Bachelor's degree or equivalent experience in a relevant discipline demonstrated through academic or professional credentials.
The role operates under a hybrid workplace policy requiring at least 25% office presence, though some positions may demand more time onsite.
Visa sponsorship is available subject to role compatibility and candidate qualification, with legal assistance provided.
Diversity and Inclusion
We encourage applications from all candidates regardless of whether every qualification is met, appreciating diverse perspectives as vital to responsible AI development.
Security Advisory
Communication regarding this position will only occur via official corporate channels and authorized recruitment partners. Beware of fraudulent solicitations requesting payments or personal banking information prior to employment commencement.
Level
Mid
Minimum education
Bachelor's Degree