F

Senior Kubernetes Platform Engineer

Firmus

Sydney, New South Wales, Australia · Full Time

Be the first to apply

Experience
Any
Salary
—
Openings
1
Posted
6 days ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

AI FactoryOS Operations 

AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system. 

AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against. 

The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability. It also builds the shared services the estate's own operation depends on, and runs them. Operating the estate every day is what shows how the platform behaves under real load and under failure, and the function works with the engineering teams that build it to turn what it finds into permanent fixes and design improvements. 

Senior Kubernetes Platform Engineer

Role Summary 

Firmus runs large-scale, state-of-the-art AI infrastructure built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation. The Senior Kubernetes Platform Engineer  runs the Kubernetes estate that every product and every tenant runs on: the platform controller layer, the virtual cluster platform tenants are provisioned onto, the Kubernetes environment baselines used across the estate, the GPU integration layer, and the automated tenant onboarding and release pipeline. 

This is a hands-on senior role with deep technical expertise. This role owns the platform lifecycle execution across the fleet: keeping clusters healthy and current across the fleet, keeping tenant workloads running through upgrade and failure, and restoring control planes to service when they degrade. Automation is a first-class part of the role, in the operational tooling, guarded remediation and fleet orchestration that make estate-scale operation possible, delivered as controlled code and reviewed by AI Infrastructure where it affects service behaviour. 

Key Responsibilities 

  • Responsible for the reliable operation, automation and continuous improvement of the multi-tenant Kubernetes platform that Firmus' products and tenants run on, spanning every site in the estate. 
  • Build the operational tooling, guarded remediation and orchestration that automate operations at fleet scale, and contribute operator and controller requirements, and code where agreed, to AI Infrastructure's platform backlog with the production evidence behind them. 
  • Execute the Kubernetes cluster lifecycle across the fleet, including provisioning, patching, upgrade and decommissioning, running the deployment and upgrade mechanisms built by AI Infrastructure through the agreed staged or canary path, and holding estate version compliance and retirement coordination. 
  • Operate and recover Kubernetes control planes carrying live tenant workload, including etcd state, certificate rotation, failed upgrades and corrupted resources. 
  • Operate the virtual cluster platform and multi-tenant isolation patterns that tenants are provisioned onto, and the automated tenant onboarding and release pipeline that lands new tenants safely and repeatably. 
  • Operate the GPU integration layer for Kubernetes (for example the NVIDIA GPU Operator), including device plugins, GPU scheduling and driver coordination. 
  • Diagnose and resolve scheduling failures, CNI and CSI faults, admission rejections and resource contention from first principles, and drive continuous improvement in cluster validation, CI/CD automation, and provisioning and testing frameworks. 
  • Run the Kubernetes baselines in production carrying the admission policy, workload identity and network policy content set by the Senior Platform Security Engineer, and hold the operational acceptance requirements those baselines have to meet before they enter production. 
  • Provide the deepest technical expertise for Kubernetes faults across the estate, diagnosing the faults that require internals-level knowledge to root cause, and driving the permanent fix to closure through AI Infrastructure, and mentor the engineers who carry frontline diagnosis, documenting operational procedures, runbooks and performance results. 
  • Lead technical recovery during major Kubernetes incidents under the incident commander, drive the changes that remove repeat causes through the problem record, and share the after-hours escalation roster for the Kubernetes estate. 

 

Skills & Experience 

Tools & software

Kubernetes required
🤖
Online · instant AI help
Broxer