Software Engineering Manager, Site Reliability Engineering (SRE), AI Foundations Data Intelligence
Sydney, New South Wales, Australia · Full Time
Be the first to apply
- Experience
- 8+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 5 hours ago
- Work mode
- In office
- Education
- Bachelor’s degree in Computer Science or related field
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Google and Commitment to Diversity
Google strives to empower and provide equitable opportunities for Aboriginal and Torres Strait Islander peoples as part of its Reconciliation Action Plan. Indigenous applicants are encouraged to apply.
Role Overview
This managerial position within the Site Reliability Engineering (SRE) team focuses on building and maintaining large-scale, distributed, fault-tolerant systems that ensure Google's critical services deliver high reliability and fast improvements. The SRE team combines software and systems engineering principles to optimize infrastructure through automation, handle scalability challenges, and maintain system performance and capacity.
Key Responsibilities
- Lead and mentor a geographically distributed team of SRE engineers, managing performance, career growth, and recruitment aligned with expanding project needs.
- Own overall reliability for F1 Query and CDQ services, participating actively in a Tier-1 on-call schedule to quickly resolve incidents and maintain high availability with a 5-minute response time SLA.
- Identify system bottlenecks and architect resilient solutions, replacing manual tasks with automation to sustainably scale infrastructure while enhancing reliability and deployment speed.
- Collaborate with core Query development teams and other SRE groups across AI Foundations and Data Intelligence to synchronize project roadmaps and spearhead reliability initiatives spanning the organization.
- Provide technical leadership through design reviews, consulting, and training in production environments, influencing engineering practices for services written in C++, Java, and Go, including capacity planning and performance optimization.
Qualifications
- Bachelor’s degree in Computer Science, related discipline, or equivalent hands-on experience.
- Minimum of 8 years experience in software development involving one or multiple programming languages.
- At least 3 years of managerial experience overseeing teams or people.
- Minimum 3 years leading projects successfully.
- At least 3 years designing, analyzing, and troubleshooting distributed systems.
- Master’s degree in Computer Science or Engineering is preferred but not mandatory.
Additional Information
F1 Query (AI Foundations) is a Google-developed, planet-scale distributed federated query engine compliant with GoogleSQL, enabling cross-source joins across core internal storage systems such as Spanner, Capacitor, and ColumnIO. This system supports over 400 production systems including Ads, Cloud, YouTube, and DeepMind platforms for low-latency analytics and batch processing.
The Technical Infrastructure team builds and maintains the architecture behind Google’s online services, manages data centers, and develops next-generation platforms ensuring optimal user experience and network reliability.
Equal Opportunity Employment
Google is an equal opportunity employer dedicated to diversity and inclusion, hiring across all identities without discrimination. Accommodation needs for applicants with disabilities can be communicated during the process.
Minimum education
Bachelor's Degree