Aziro

Senior Staff Software Engineer, Site Reliability & Platform Automation

Aziro

Bengaluru, Karnataka, India · Full Time

Be the first to apply

Experience
12+ yrs
Salary
—
Openings
1
Posted
3 days ago
Work mode
In office
Education
B.Tech / B.E.
Eligibility
Candidates with a Bachelor of Technology (B.Tech) or Bachelor of Engineering (B.E.) in any specialization are eligible to apply.
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Aziro

Aziro is a pioneering company specializing in IT services and industry-specific automation software. Their unique approach, known as the Aziro Way, fosters innovation and collaboration, delivering exceptional value in storage, servers, disaster recovery/business continuity, and virtualization at competitive costs. Alongside its services, Aziro offers cutting-edge domain-specific software solutions for test automation, aiming to be a transformational force in the industry by delivering value-added software and optimized business processes.

Job Overview

This position as Senior Staff Software Engineer in Site Reliability & Platform Automation involves joining the SaaS Platform Engineering team in Bangalore, India. Reporting to the Senior Manager of Site Reliability & Platform Engineering, the role requires setting strategic technical direction to enhance reliability measurement, automate operational toil, conduct resilience testing, and enable self-service functionalities across Infoblox's global cloud networking and SaaS platforms. Collaboration with cross-functional teams such as DevOps, CloudOps, product engineering, architecture, security, and product management is fundamental.

Key Responsibilities

  • Define and unify technical strategy encompassing reliability metrics, automation of manual tasks, resilience verification, and enabling self-service capabilities ensuring integrated platform architecture.
  • Lead design efforts related to high availability, Service Level Objectives/Indicators (SLO/SLI), error budgets, chaos and disaster recovery testing, change management governance, incident resolution, observability, and tuning application performance and capacity.
  • Create and manage golden paths, Terraform modules, and Kyverno policies to empower product teams to independently provision and manage services via safe guardrails.
  • Develop measurement pipelines and evidence repositories to quantify reliability, service health, developer experience, operational efficiency, and platform usage.
  • Construct automation frameworks for tasks such as CVE remediation, infrastructure upgrades, third-party provider testing, and everyday operational procedures across multiple production environments.
  • Collaborate with North America and India DevOps and CloudOps teams to analyze and replace manual workflows with dependable, maintainable software automation.
  • Evaluate whether to build custom solutions or adopt existing open-source, commercial, or internal tools to minimize unnecessary platform maintenance overhead.
  • Author and review production-grade software in languages such as Go, Python, Rust, and Java, including comprehensive testing, documentation, instrumentation, and operational safeguards.
  • Design resilient, secure, and observable cloud environments on AWS and GCP, ensuring compliance with standards like FedRAMP, SOC 2, or ISO when required.
  • Apply AI-driven engineering tools responsibly to boost productivity, automate workflows, analyze telemetry, generate content, and support incident decision-making while maintaining necessary human oversight, security, and accountability.

Candidate Profile

  • Extensive software engineering experience totaling more than 12 years, with significant expertise in infrastructure, platform development, developer tooling, distributed systems, or reliability engineering.
  • Proven track record in creating and refining internal platforms, tools, or reliability systems actively used by engineering teams in production.
  • Strong specialization in at least two domains among distributed systems, Kubernetes/container platforms, infrastructure as code, CI/CD pipelines, and observability/telemetry systems, with working proficiency across others.
  • Robust software development skills in languages including Go, Python, Rust, or Java, with capabilities in testing, debugging, performance tuning, and operational responsibility.
  • Hands-on experience designing and managing cloud platforms on AWS and/or GCP, including proficiency with Terraform and Kubernetes in infrastructure automation.
  • Familiarity with establishing or scaling SLIs/SLOs, error budgets, observability frameworks, incident management processes, high availability strategies, disaster recovery, and resilience verification.
  • Ability to automate repetitive manual tasks effectively, demonstrably improving toil, reliability, scalability, change management safety, or developer experience.
  • Adopts a platform-as-product approach, focusing on internal engineers as customers and tracking key adoption and satisfaction metrics.
  • Experience using AI-based engineering or operational tools responsibly for productivity enhancement, analytics, content creation, and decision support.
  • Excellent interpersonal skills including influencing senior stakeholders, mentoring peers, articulating complex ideas, and fostering adoption across organizational teams. A bachelor's or master's degree in computer science or related fields is preferred.

Preferred Qualifications

  • Proven experience managing multi-tenant, multi-region, or multi-cloud Kubernetes and cloud infrastructures at significant scale.
  • Knowledge of regulated environments, with exposure to compliance standards like FedRAMP, SOC 2, or ISO.
  • Familiarity with observability tools such as Prometheus, Grafana, OpenTelemetry, Loki, ELK stack, Datadog, or equivalents.
  • Experience with policy-as-code frameworks including Kyverno, OPA, or Gatekeeper.
  • Competence in CI/CD systems like Jenkins, GitHub Actions, GitOps, or Argo.
  • Background in chaos engineering, fault injection, disaster recovery exercises, failover strategies, or resilience testing at production scale.
  • Knowledge of large language model (LLM)-based or agentic tooling applied to operations with appropriate safety controls.

Eligibility

Applicants must possess a Bachelor of Technology (B.Tech) or Bachelor of Engineering (B.E.) degree in any field of specialization.

Minimum education

Bachelor's Degree

Tools & software

Kubernetes required Terraform required

How they work

Communication Teamwork & Collaboration Leadership Creativity

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer