I
Senior Infrastructure & Reliability Engineer - AI-DNA
Remote · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Overview
We are developing an AI-centric Infrastructure & Reliability division where autonomous AI agents take charge of incident investigations, hypothesis validation, root cause analysis generation, and secure remediation of production issues.
This role of Senior Infrastructure & Reliability Engineer involves leading ownership of production reliability while enhancing AI agents that handle incident responses, operational workflows, and infrastructure management.
Key Responsibilities
- Maintain and assure reliability in production environments by promptly addressing customer-impacting incidents.
- Design, develop, and refine AI-driven agents for incident triage, change validation, root cause analysis (RCA) creation, and automated remediation processes.
- Conduct production deployments and infrastructure modifications adhering to rigorous operational protocols.
- Develop production-quality code, automation scripts, operational runbooks, and AI workflow integrations aimed at reducing repetitive manual tasks.
- Drive continuous improvements in platform uptime, operational effectiveness, and overall user experience.
Qualifications and Skills
- Minimum of five years' experience operating enterprise SaaS platforms focusing on Infrastructure, Platform Engineering, DevOps, Cloud Operations, or Site Reliability Engineering.
- Proven expertise managing AWS-based production environments, including large-scale cloud infrastructures and efficient incident handling.
- Hands-on engineering aptitude with a strong inclination for diagnosing complex production system issues and implementing automation solutions.
- Familiarity with AI engineering tools like Claude Code, Codex, Cursor, Warp, or equivalent technologies.
- Strong enthusiasm for pioneering AI-driven operational automation and autonomous infrastructure management.
- Excellent command of written and spoken English for clear communication.
Benefits and Work Culture
- Opportunity to contribute to one of the pioneering AI-first Infrastructure & Reliability teams in the industry, focusing on building AI systems that minimize incident tickets rather than just resolving them.
- AI-centric work environment encouraging unrestricted use of AI tools and computational resources.
- Fully remote employment allowing flexible work arrangements.
- Diverse, global team collaboration.
- Engagement with enterprise-grade SaaS platforms.
Skills
Tools & software
Amazon Web Services AWS
required
How they work
Communication
Problem Solving
Attention to Detail
Work Ethic
Languages
Servicenow