I

Senior DevOps Engineer

IgniteTech

Remote · Full Time

Be the first to apply

Experience
5+ yrs
Salary
—
Openings
1
Posted
4 days ago
Work mode
Work from home
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

We operate the community and social engagement platform behind some of the world’s leading brands. Our service demands exceptional reliability because even a brief outage can jeopardize customer trust of Fortune 100 companies. This responsibility shapes our engineering philosophy: every manual intervention triggers automation to prevent recurrence.

Key Responsibilities

  • Take full operational ownership during your shifts, acting as the immediate responder when issues arise. Manage triage, diagnosis, mitigation, restoration, and escalate when necessary. Ensuring uptime is your personal priority.
  • Develop and refine autonomous agents and workflows that minimize manual operational effort—ranging from alert triage, deployment validations, to self-healing actions and post-incident reviews. Continuously enhance agent abilities beyond initial prompts.
  • Safeguard all production changes by implementing quality gates and proven rollback mechanisms. Respond swiftly to any deviation in telemetry by reverting changes.
  • Conduct thorough root cause investigations that distinguish symptoms from real underlying issues, followed by implementing lasting fixes to prevent recurrence and tracking their completion.
  • Transform every manual fix into scalable automation through new agent rules, guardrails, or runbooks. Focus on expanding autonomous workload to reduce manual intervention.
  • Document decisions, procedures, and operational insights to preserve institutional knowledge accessible to both autonomous agents and team members, particularly crucial for a globally distributed asynchronous environment.

Required Qualifications and Skills

  • More than 5 years of experience in roles such as Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure on large-scale SaaS platforms, including hands-on pager duty and incident management.
  • Extensive expertise with AWS, including multi-AZ and multi-account setups, infrastructure automation, rigorous gated change processes, and rollback discipline. Proven experience managing and recovering from significant outages.
  • Exceptional self-motivation and initiative to prioritize and address the most impactful operational gaps independently. Ability to challenge and improve standards rather than circumvent them.
  • Practical experience working with AI-driven autonomous operational tools, delegating work to agents, evaluating outputs critically, and enhancing agent capabilities. Familiarity with tools such as Claude Code, Codex, Warp, or related custom frameworks, and enthusiasm for testing cutting-edge AI models.
  • Possession of an AWS Solutions Architect – Associate certification or equivalent real-world proficiency.
  • Excellent English communication skills, able to maintain clarity during critical incidents and compose detailed post-mortems.
  • Commitment to shift-based coverage involving on-call rotations within assigned time zones.
  • Must reside in an OFAC-compliant country.

Preferred Attributes

  • Contributions to agentic operations, AIOps, or intelligent automation in the form of developed tools, open-source projects, published technical articles, or conference presentations.
  • Experience with multi-tenant B2B SaaS platforms such as community engagement tools, customer experience products, or observability systems.
  • Familiarity with observability and alerting technologies including Grafana, Prometheus, Datadog, PagerDuty, and OpsGenie; additional experience with Azure cloud is a plus.
  • Demonstrated deep focus on solving complex technical challenges personally and systematically.

What You Will Gain

You will play a key role in shaping a novel operational practice—agentic DevOps at enterprise scale—supporting platforms trusted daily by Fortune 100 companies. The skills even the industry is still defining, including agent engineering, incident management patterns, and system design judgment, will be part of your expertise.

Work Environment

  • Operate within an environment combining the responsibilities linked to enterprise SLAs with the rapid pace and adaptability of a startup.
  • Access to unlimited tooling investments, including advanced AI models, compute resources, and new software as needed.
  • Join a fully remote, globally distributed team.

Level

Senior

Tools & software

Amazon Web Services AWS required

How they work

Organisation required

Languages

Servicenow

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer