Z

DevOps Lead - Platform & Infrastructure Engineering

Zig by ComfortDelGro

Remote · Full Time

Be the first to apply

Experience
7+ yrs
Salary
—
Openings
1
Posted
6 days ago
Work mode
Work from home
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Job Overview

We are seeking a seasoned DevOps Lead to oversee and evolve our AWS infrastructure and platform engineering initiatives. This role involves enhancing existing Terraform codebases, managing Kubernetes clusters, optimizing cloud expenditures, and guiding the transition toward modern platform practices to support scalable organizational growth.

Key Responsibilities

  • Maintain and improve AWS infrastructure and Terraform codebases with robust remote state management and modular architecture.
  • Proactively manage infrastructure upgrades including AWS service deprecations, EKS platform updates, and Terraform provider enhancements to reduce technical debt and security vulnerabilities.
  • Conduct thorough cost analysis to optimize cloud spending related to compute resources, storage lifecycle, and cross-zone data transfers without sacrificing system resilience.
  • Design and implement next-generation platform engineering strategies such as policy-as-code and self-service infrastructure portals.
  • Ensure high availability and zero downtime for Amazon EKS production clusters, refining upgrades, node group management, and network configurations.
  • Oversee ArgoCD deployments focusing on configuration drift control, synchronization policies, and multi-cluster consistency.
  • Optimize container workload performance by adjusting resource allocations, autoscaling parameters, and reinforcing cluster security perimeters.
  • Manage and scale the self-hosted Bitbucket Runners, guaranteeing rapid and cost-effective build processes with high availability.
  • Maintain and enhance Bitbucket Pipelines to expedite developer feedback loops, integrating automation for compliance and security.
  • Develop and sustain comprehensive observability frameworks covering metrics, logging, and tracing for microservices.
  • Monitor service level agreements, reduce false alerts, and act as a senior escalation point for critical incidents, orchestrating in-depth post-incident analysis and preventive measures.
  • Lead and mentor DevOps engineers, establishing stringent standards for infrastructure code quality, documentation, and operational excellence.
  • Collaborate with software leads to translate operational challenges into strategic platform improvements.
  • Utilize Generative AI and automation to streamline workflows, minimize manual effort, enhance decision-making, and build strategic consulting capabilities within the team.
  • Perform additional duties as necessary.

Candidate Requirements

  • Preferably 7 or more years in DevOps, Cloud Operations, or SRE roles, with at least 2 years managing engineering teams or complex infrastructures.
  • Extensive hands-on experience managing core AWS services such as VPC, EC2, IAM, S3, RDS, Route53, and CloudFront in production-scale environments.
  • Expertise in maintaining and evolving large Terraform codebases that utilize multi-workspace and remote state configurations.
  • Practical knowledge of managing Kubernetes clusters (EKS) in production settings and employing ArgoCD for continuous delivery across multiple environments.
  • Demonstrated ability to administer, troubleshoot, and optimize self-hosted CI/CD runners (specifically Bitbucket Runners) and deployment pipelines.
  • Proficient scripting skills in Python, Go, or Bash for automating operational procedures and platform tasks.
  • Experience using Kubernetes autoscalers like Karpenter to dynamically provision cost-efficient EKS nodes.
  • Familiarity with integrating tools for automated compliance, vulnerability scanning, and policy enforcement such as Trivy, Checkov, and OPA.
  • Certifications like Certified Kubernetes Administrator (CKA) or AWS Certified DevOps Engineer – Professional are advantageous.

Tools & software

Kubernetes · 2 to 5 years required Amazon Web Services AWS required Argocd · 1 to 2 years required Terraform · 5 to 8 years required

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer