- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 13 hours ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
Join our team as a Senior DevOps & Autonomous Systems Engineer, specializing in maintaining and improving the reliability of a community and social engagement platform used by Fortune 100 companies. The core mindset is to minimize human intervention by developing automation that prevents operational issues from recurring.
Key Responsibilities
- Take full ownership of your shift, acting as the first responder during system incidents to diagnose, mitigate, restore, or escalate issues promptly, with uptime as a personal commitment.
- Develop autonomous workflows and agents to reduce manual operational tasks, including alert pre-triage, deployment checks, self-healing sequences, and post-incident reviews, continuously enhancing their capabilities.
- Protect production environments by ensuring every deployment, configuration, or cost-related action passes strict quality checks with reliable rollback mechanisms, withdrawing changes immediately upon detecting anomalies.
- Conduct thorough root cause analysis to identify and fix underlying issues, delivering permanent preventative solutions rather than temporary patches.
- Transform manual fixes into durable improvements by creating new agent protocols, guardrails, or runbooks, expanding the scope of automation and reducing reliance on manual interventions.
- Document decisions, procedures, and operational context effectively to preserve institutional knowledge accessible by both humans and autonomous systems, essential for a globally distributed asynchronous team.
Required Qualifications
- At least 5 years of hands-on experience in production operations roles such as Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure within large-scale SaaS environments, with proven incident management expertise.
- Deep expertise in AWS architecture at scale, including multi-availability zone and multi-account setups, infrastructure automation, controlled change management, and rollback procedures, with firsthand experience managing significant outages.
- Strong self-motivation and independence to proactively identify and address critical operational gaps without waiting for instructions, and the willingness to improve or challenge existing standards.
- Practical experience working with AI-driven operational tools, delegating tasks to automated agents, critically evaluating their outputs, and enhancing underlying agent frameworks; familiarity with technologies like Claude Code, Codex, Warp, or similar custom frameworks is preferred.
- Credentials such as AWS Solutions Architect – Associate or equivalent substantial experience in AWS infrastructure management.
- Excellent communication skills in English, both during high-pressure incidents and detailed post-incident analyses.
- Commitment to participating in on-call shifts within a designated time zone, as shift coverage is critical to this role.
- Residency in a country compliant with OFAC regulations.
Preferred Attributes
- Contributions to agentic operations, AIOps, or intelligent automation through tools developed, open-source contributions, published technical content, or conference presentations.
- Experience working with multi-tenant B2B SaaS platforms such as community platforms, social media tools, customer experience products, or monitoring systems.
- Proficiency with observability platforms and alerting tools like Grafana, Prometheus, Datadog, PagerDuty, and OpsGenie; familiarity with Azure is a plus in addition to AWS expertise.
- A demonstrated deep-focus passion for intricate technical challenges that showcase analytical thinking and problem-solving skills beyond deliverables.
Benefits and Work Environment
- Operate at the intersection of enterprise-level reliability and startup agility, managing SLAs for Fortune 100 clients while working in rapid iteration cycles and continuously refining processes.
- Access to unrestricted tooling and resources to innovate, including options for larger AI models, more computing power, or new tools as needed.
- Work fully remotely within a globally distributed team, emphasizing asynchronous collaboration and documentation-driven institutional memory.
Level
Senior
Skills
How they work
Communication
Problem Solving
Attention to Detail
Independence
Accountability
Languages
English