I

Senior DevOps Engineer

IgniteTech

Remote · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
13 hours ago
Work mode
Work from home
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

Join our team as a Senior DevOps & Autonomous Systems Engineer, specializing in maintaining and improving the reliability of a community and social engagement platform used by Fortune 100 companies. The core mindset is to minimize human intervention by developing automation that prevents operational issues from recurring.

Key Responsibilities

  • Take full ownership of your shift, acting as the first responder during system incidents to diagnose, mitigate, restore, or escalate issues promptly, with uptime as a personal commitment.
  • Develop autonomous workflows and agents to reduce manual operational tasks, including alert pre-triage, deployment checks, self-healing sequences, and post-incident reviews, continuously enhancing their capabilities.
  • Protect production environments by ensuring every deployment, configuration, or cost-related action passes strict quality checks with reliable rollback mechanisms, withdrawing changes immediately upon detecting anomalies.
  • Conduct thorough root cause analysis to identify and fix underlying issues, delivering permanent preventative solutions rather than temporary patches.
  • Transform manual fixes into durable improvements by creating new agent protocols, guardrails, or runbooks, expanding the scope of automation and reducing reliance on manual interventions.
  • Document decisions, procedures, and operational context effectively to preserve institutional knowledge accessible by both humans and autonomous systems, essential for a globally distributed asynchronous team.

Required Qualifications

  • At least 5 years of hands-on experience in production operations roles such as Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure within large-scale SaaS environments, with proven incident management expertise.
  • Deep expertise in AWS architecture at scale, including multi-availability zone and multi-account setups, infrastructure automation, controlled change management, and rollback procedures, with firsthand experience managing significant outages.
  • Strong self-motivation and independence to proactively identify and address critical operational gaps without waiting for instructions, and the willingness to improve or challenge existing standards.
  • Practical experience working with AI-driven operational tools, delegating tasks to automated agents, critically evaluating their outputs, and enhancing underlying agent frameworks; familiarity with technologies like Claude Code, Codex, Warp, or similar custom frameworks is preferred.
  • Credentials such as AWS Solutions Architect – Associate or equivalent substantial experience in AWS infrastructure management.
  • Excellent communication skills in English, both during high-pressure incidents and detailed post-incident analyses.
  • Commitment to participating in on-call shifts within a designated time zone, as shift coverage is critical to this role.
  • Residency in a country compliant with OFAC regulations.

Preferred Attributes

  • Contributions to agentic operations, AIOps, or intelligent automation through tools developed, open-source contributions, published technical content, or conference presentations.
  • Experience working with multi-tenant B2B SaaS platforms such as community platforms, social media tools, customer experience products, or monitoring systems.
  • Proficiency with observability platforms and alerting tools like Grafana, Prometheus, Datadog, PagerDuty, and OpsGenie; familiarity with Azure is a plus in addition to AWS expertise.
  • A demonstrated deep-focus passion for intricate technical challenges that showcase analytical thinking and problem-solving skills beyond deliverables.

Benefits and Work Environment

  • Operate at the intersection of enterprise-level reliability and startup agility, managing SLAs for Fortune 100 clients while working in rapid iteration cycles and continuously refining processes.
  • Access to unrestricted tooling and resources to innovate, including options for larger AI models, more computing power, or new tools as needed.
  • Work fully remotely within a globally distributed team, emphasizing asynchronous collaboration and documentation-driven institutional memory.

Level

Senior

How they work

Communication Problem Solving Attention to Detail Independence Accountability

Languages

English

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help