Wipro

DevOps Site Reliability Engineer (Kubernetes Multi-Cloud)

Wipro

Singapore · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
30 seconds ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

As a DevOps Site Reliability Engineer specialized in Kubernetes, you will serve as the initial point of contact for troubleshooting and resolving Kubernetes cluster issues across AWS, Onprem AliCloud, and GCP environments. This role demands complete ownership of support tickets from their creation until resolution through a Slack-based channel, escalating to platform engineering only for confirmed defects or capacity changes. Based in Singapore, you will primarily manage the APAC Kubernetes cluster landscape.

Key Responsibilities

  • Manage Kubernetes cluster lifecycle, including creation, updates, deletions, node pool handling, and recovery from stuck states.
  • Perform Kubernetes version upgrades and diagnose associated failures.
  • Handle RBAC configurations and related impacts during node rotations.
  • Troubleshoot kubeconfig, kubectl issues, API server performance and request throttling.
  • Operate Kubernetes clusters on multiple cloud providers including AWS (EKS, Karpenter, Route 53), AliCloud (ACK), and GCP (GKE/GCE), ensuring region-specific version compliance and vulnerability remediation.
  • Manage platform tooling such as Kubernetes manifests, YAML validation, deployment monitoring, and CLI usage.
  • Oversee observability through Prometheus, kube-state-metrics, and Splunk forwarding.
  • Implement and maintain add-ons such as Spinnaker, KEDA, Compass compliance, and backup services.
  • Resolve networking issues covering DNS resolution, cluster connectivity, load balancer reconciliation, and subnet annotations.
  • Conduct outage investigations and author formal Root Cause Analyses.

Required Qualifications

  • Minimum five years of experience in Kubernetes operations, cloud infrastructure support, SRE, or cloud operations engineering with multi-cloud exposure.
  • Extensive hands-on experience operating managed Kubernetes environments, notably EKS or GKE, with competency in a second cloud provider.
  • Strong diagnostic skills related to cluster health, pod/node state, storage and network interface issues, cluster events, and control plane behavior under load.
  • Solid understanding of identity and access management concepts across multiple cloud providers, including role delegation, workload identities, service accounts, and integrating IAM with Kubernetes RBAC.
  • Proficient in enterprise networking fundamentals: routing, DNS, TLS, proxies, CIDR planning, firewall/ACL models, and their implementation in ingress and load balancer controllers.
  • Experience with scripting in Python or Bash and adeptness using kubectl and cloud provider command-line tools.
  • Capability to independently manage support tickets with clear communication and documented resolutions.
  • Excellent written English skills to support asynchronous communication efficiently.
  • Legal authorization to work in Singapore.

Preferred Qualifications

  • Certifications such as CKA, CKAD, AWS Solutions Architect (Associate or Professional), Google Professional Cloud Architect, or Alibaba Cloud ACP.
  • Hands-on experience with AliCloud or rapid adaptability to new cloud providers, especially AliCloud ACK.
  • Background in supporting internal developer platforms within large organizations.
  • Knowledge of hybrid connectivity between corporate networks and public clouds including on-premise Kubernetes setups.
  • Experience with infrastructure-as-code tools like Terraform, CloudFormation and templating frameworks such as Helm or Kustomize.
  • Familiarity with observability platforms including Splunk, Prometheus, Grafana, Datadog, CloudWatch, and Cloud Logging.
  • Experience participating in follow-the-sun or multi-region support models.
  • Understanding of China-region cloud operational challenges and regulatory requirements.
  • Proven track record of automating recurring escalations to reduce ticket volume through runbook creation and automation.

Tools & software

AWS Kubernetes · 5 to 8 years required Prometheus required Splunk Enterprise required Terraform required

How they work

Communication Problem Solving
🤖
Online · instant AI help
Broxer