Senior Staff Software Engineer, Site Reliability & Platform Automation
Bengaluru, Karnataka, India · Full Time
Be the first to apply
- Experience
- 12+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 3 days ago
- Work mode
- In office
- Education
- B.Tech / B.E.
- Eligibility
- Candidates with a Bachelor of Technology (B.Tech) or Bachelor of Engineering (B.E.) in any specialization are eligible to apply.
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Aziro
Aziro is a pioneering company specializing in IT services and industry-specific automation software. Their unique approach, known as the Aziro Way, fosters innovation and collaboration, delivering exceptional value in storage, servers, disaster recovery/business continuity, and virtualization at competitive costs. Alongside its services, Aziro offers cutting-edge domain-specific software solutions for test automation, aiming to be a transformational force in the industry by delivering value-added software and optimized business processes.
Job Overview
This position as Senior Staff Software Engineer in Site Reliability & Platform Automation involves joining the SaaS Platform Engineering team in Bangalore, India. Reporting to the Senior Manager of Site Reliability & Platform Engineering, the role requires setting strategic technical direction to enhance reliability measurement, automate operational toil, conduct resilience testing, and enable self-service functionalities across Infoblox's global cloud networking and SaaS platforms. Collaboration with cross-functional teams such as DevOps, CloudOps, product engineering, architecture, security, and product management is fundamental.
Key Responsibilities
- Define and unify technical strategy encompassing reliability metrics, automation of manual tasks, resilience verification, and enabling self-service capabilities ensuring integrated platform architecture.
- Lead design efforts related to high availability, Service Level Objectives/Indicators (SLO/SLI), error budgets, chaos and disaster recovery testing, change management governance, incident resolution, observability, and tuning application performance and capacity.
- Create and manage golden paths, Terraform modules, and Kyverno policies to empower product teams to independently provision and manage services via safe guardrails.
- Develop measurement pipelines and evidence repositories to quantify reliability, service health, developer experience, operational efficiency, and platform usage.
- Construct automation frameworks for tasks such as CVE remediation, infrastructure upgrades, third-party provider testing, and everyday operational procedures across multiple production environments.
- Collaborate with North America and India DevOps and CloudOps teams to analyze and replace manual workflows with dependable, maintainable software automation.
- Evaluate whether to build custom solutions or adopt existing open-source, commercial, or internal tools to minimize unnecessary platform maintenance overhead.
- Author and review production-grade software in languages such as Go, Python, Rust, and Java, including comprehensive testing, documentation, instrumentation, and operational safeguards.
- Design resilient, secure, and observable cloud environments on AWS and GCP, ensuring compliance with standards like FedRAMP, SOC 2, or ISO when required.
- Apply AI-driven engineering tools responsibly to boost productivity, automate workflows, analyze telemetry, generate content, and support incident decision-making while maintaining necessary human oversight, security, and accountability.
Candidate Profile
- Extensive software engineering experience totaling more than 12 years, with significant expertise in infrastructure, platform development, developer tooling, distributed systems, or reliability engineering.
- Proven track record in creating and refining internal platforms, tools, or reliability systems actively used by engineering teams in production.
- Strong specialization in at least two domains among distributed systems, Kubernetes/container platforms, infrastructure as code, CI/CD pipelines, and observability/telemetry systems, with working proficiency across others.
- Robust software development skills in languages including Go, Python, Rust, or Java, with capabilities in testing, debugging, performance tuning, and operational responsibility.
- Hands-on experience designing and managing cloud platforms on AWS and/or GCP, including proficiency with Terraform and Kubernetes in infrastructure automation.
- Familiarity with establishing or scaling SLIs/SLOs, error budgets, observability frameworks, incident management processes, high availability strategies, disaster recovery, and resilience verification.
- Ability to automate repetitive manual tasks effectively, demonstrably improving toil, reliability, scalability, change management safety, or developer experience.
- Adopts a platform-as-product approach, focusing on internal engineers as customers and tracking key adoption and satisfaction metrics.
- Experience using AI-based engineering or operational tools responsibly for productivity enhancement, analytics, content creation, and decision support.
- Excellent interpersonal skills including influencing senior stakeholders, mentoring peers, articulating complex ideas, and fostering adoption across organizational teams. A bachelor's or master's degree in computer science or related fields is preferred.
Preferred Qualifications
- Proven experience managing multi-tenant, multi-region, or multi-cloud Kubernetes and cloud infrastructures at significant scale.
- Knowledge of regulated environments, with exposure to compliance standards like FedRAMP, SOC 2, or ISO.
- Familiarity with observability tools such as Prometheus, Grafana, OpenTelemetry, Loki, ELK stack, Datadog, or equivalents.
- Experience with policy-as-code frameworks including Kyverno, OPA, or Gatekeeper.
- Competence in CI/CD systems like Jenkins, GitHub Actions, GitOps, or Argo.
- Background in chaos engineering, fault injection, disaster recovery exercises, failover strategies, or resilience testing at production scale.
- Knowledge of large language model (LLM)-based or agentic tooling applied to operations with appropriate safety controls.
Eligibility
Applicants must possess a Bachelor of Technology (B.Tech) or Bachelor of Engineering (B.E.) degree in any field of specialization.
Minimum education
Bachelor's Degree