- Experience
- 8–12 yrs
- Salary
- —
- Openings
- 1
- Posted
- 6 days ago
- Work mode
- In office
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Overview
We are seeking a seasoned DevOps Engineer with 8 to 12 years of deep expertise in cloud engineering and DevOps practices. This role requires proficient hands-on experience managing AWS environments, container orchestration with Kubernetes, and infrastructure automation. The ideal candidate will be capable of designing scalable, secure, and cost-efficient cloud infrastructure while driving observability and security requirements.
Key Responsibilities
- Architect and maintain scalable, secure, and cost-optimized infrastructure within AWS cloud.
- Develop and manage Kubernetes workloads, utilizing EKS, ECS, Docker, Helm, and service mesh technologies.
- Create and sustain Infrastructure as Code (IaC) implementations primarily via Terraform.
- Construct and enable CI/CD pipelines with tools such as Azure Pipelines and GitHub Actions.
- Administer AWS serverless and event-driven resources, including Lambda, API Gateway, SQS/SNS, and EventBridge.
- Establish and maintain observability stacks using OpenTelemetry, Prometheus, Grafana, AWS CloudWatch, and ELK.
- Define and monitor service reliability metrics such as SLIs, SLOs, and SLAs.
- Lead initiatives around cloud security, including identity and access management (IAM), key management, secrets management, compliance, and vulnerability scanning.
- Plan and implement backup and disaster recovery strategies aligned with required recovery time objective (RTO) and recovery point objective (RPO).
- Manage production incident response activities, including troubleshooting, root cause analysis, and post-mortem documentation.
- Continuously enhance infrastructure reliability, automation levels, performance efficiencies, and cloud cost management.
- Leverage AI and Copilot tools to assist in writing and reviewing Terraform scripts, CI/CD pipelines, and automation scripts.
- Utilize AIOps capabilities for anomaly detection, log summarization, and AI-supported incident root cause analysis.
- Support infrastructure tailored for AI workloads such as GPU and inference nodes, vector databases, and model-serving platforms.
Qualifications and Experience
- 8 to 12 years of relevant experience in Cloud/DevOps engineering roles.
- Hands-on proficiency with core AWS services and platform features.
- In-depth experience managing Kubernetes with EKS, ECS, Docker containers, Helm charts, ingress controllers, and service mesh technologies.
- Strong capability in Terraform for IaC; familiarity with AWS CloudFormation or CDK is desirable.
- Proven experience building CI/CD pipelines using Azure Pipelines and GitHub Actions.
- Advanced scripting skills in Python, Bash, or PowerShell for automation purposes.
- Competence using observability and monitoring tools such as OpenTelemetry, Prometheus, Grafana, AWS CloudWatch, and ELK stack.
- Thorough understanding of AWS security services including IAM, KMS, secrets management, and security assessments.
- Comfortable using Git version control and working within Agile/Scrum team processes.
- Exemplary incident management, troubleshooting, root cause analysis, and conducting post-incident reviews.
- Knowledge of designing systems for high availability, disaster recovery, and optimizing to meet RTO and RPO goals while managing costs.
Additional Details
- Work Location: Pune, Maharashtra — Onsite position requiring office presence.
- Interview Process: Includes multiple stages — technical screening, technical interview, system design interview, hiring manager discussion, and HR round.
Tools & software
Microsoft Azure
Microsoft Azure
required
How they work
Teamwork & Collaboration
Problem Solving