Z
DevOps Lead - Platform & Infrastructure Engineering
Remote · Full Time
Be the first to apply
- Experience
- 7+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 6 days ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Job Overview
We are seeking a seasoned DevOps Lead to oversee and evolve our AWS infrastructure and platform engineering initiatives. This role involves enhancing existing Terraform codebases, managing Kubernetes clusters, optimizing cloud expenditures, and guiding the transition toward modern platform practices to support scalable organizational growth.
Key Responsibilities
- Maintain and improve AWS infrastructure and Terraform codebases with robust remote state management and modular architecture.
- Proactively manage infrastructure upgrades including AWS service deprecations, EKS platform updates, and Terraform provider enhancements to reduce technical debt and security vulnerabilities.
- Conduct thorough cost analysis to optimize cloud spending related to compute resources, storage lifecycle, and cross-zone data transfers without sacrificing system resilience.
- Design and implement next-generation platform engineering strategies such as policy-as-code and self-service infrastructure portals.
- Ensure high availability and zero downtime for Amazon EKS production clusters, refining upgrades, node group management, and network configurations.
- Oversee ArgoCD deployments focusing on configuration drift control, synchronization policies, and multi-cluster consistency.
- Optimize container workload performance by adjusting resource allocations, autoscaling parameters, and reinforcing cluster security perimeters.
- Manage and scale the self-hosted Bitbucket Runners, guaranteeing rapid and cost-effective build processes with high availability.
- Maintain and enhance Bitbucket Pipelines to expedite developer feedback loops, integrating automation for compliance and security.
- Develop and sustain comprehensive observability frameworks covering metrics, logging, and tracing for microservices.
- Monitor service level agreements, reduce false alerts, and act as a senior escalation point for critical incidents, orchestrating in-depth post-incident analysis and preventive measures.
- Lead and mentor DevOps engineers, establishing stringent standards for infrastructure code quality, documentation, and operational excellence.
- Collaborate with software leads to translate operational challenges into strategic platform improvements.
- Utilize Generative AI and automation to streamline workflows, minimize manual effort, enhance decision-making, and build strategic consulting capabilities within the team.
- Perform additional duties as necessary.
Candidate Requirements
- Preferably 7 or more years in DevOps, Cloud Operations, or SRE roles, with at least 2 years managing engineering teams or complex infrastructures.
- Extensive hands-on experience managing core AWS services such as VPC, EC2, IAM, S3, RDS, Route53, and CloudFront in production-scale environments.
- Expertise in maintaining and evolving large Terraform codebases that utilize multi-workspace and remote state configurations.
- Practical knowledge of managing Kubernetes clusters (EKS) in production settings and employing ArgoCD for continuous delivery across multiple environments.
- Demonstrated ability to administer, troubleshoot, and optimize self-hosted CI/CD runners (specifically Bitbucket Runners) and deployment pipelines.
- Proficient scripting skills in Python, Go, or Bash for automating operational procedures and platform tasks.
- Experience using Kubernetes autoscalers like Karpenter to dynamically provision cost-efficient EKS nodes.
- Familiarity with integrating tools for automated compliance, vulnerability scanning, and policy enforcement such as Trivy, Checkov, and OPA.
- Certifications like Certified Kubernetes Administrator (CKA) or AWS Certified DevOps Engineer – Professional are advantageous.
Skills
Tools & software
Kubernetes
· 2 to 5 years required
Amazon Web Services AWS
required
Argocd
· 1 to 2 years required
Terraform
· 5 to 8 years required