Senior Release Engineer
Abu Dhabi Emirate, United Arab Emirates · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Company Overview
Open Innovation AI is a technology firm dedicated to creating advanced solutions that streamline AI workload management. Their main product, the Open Innovation Cluster Manager (OICM), efficiently orchestrates AI jobs across varied hardware infrastructures, supporting multiple GPUs and accelerators. The platform is designed for seamless scalability and integration, aimed at enterprises seeking optimized and simplified AI deployment methods to reduce costs and speed up value realization.
Role Overview
The Senior Release Engineer is responsible for managing the entire build pipeline for the platform, ensuring delivery of validated, deployable software. This includes creating hardened OS images, installer ISOs, container and Helm packages, offline mirrors for packages and registries, and the orchestration environment using Ansible. The role also leads development of an automated release-validation framework to provision full clusters and verify in-place upgrades before deployment, significantly mitigating deployment risks.
Key Responsibilities
- Manage the complete air-gapped build pipeline including hardened OS images with Packer, installer ISOs, container images, Helm charts, offline package mirrors, and Ansible environments.
- Develop and maintain the release-validation gate that provisions clusters from bare metal and executes upgrade tests for every release candidate.
- Produce versioned, self-contained offline release artifacts that are customer-ready.
- Maintain continuous integration systems for building, testing, and validating platform changes.
- Handle security patching and CVE responses including rapid rebuild and validation of affected components.
- Administer offline registries and artifact servers supporting builds and deployed clusters.
- Ensure builds are reproducible and deployments are strictly self-contained without external dependencies.
- Document all processes thoroughly via architecture decision records, runbooks, and release guides for seamless knowledge transfer.
- Collaborate with infrastructure leadership to reinforce and safeguard the release workflow against regressions and unsafe upgrades.
Required Qualifications and Experience
- Proficiency with Bash and Python scripting focused on creating reusable, maintainable automation tools.
- Strong software architecture skills for designing robust, maintainable build and test systems.
- Expertise in Linux system administration, including OS packaging, image building (preferably Packer), service management, and troubleshooting.
- Experience building and managing Docker/OCI container images and Helm chart packaging and templating.
- Comprehensive understanding of CI/CD pipeline ownership, versioning, and producing reproducible release artifacts.
- Ability to develop automated end-to-end integration tests orchestrating real infrastructure provisioning and validation.
- Hands-on knowledge of Ansible for infrastructure as code.
- Familiarity with Kubernetes installation and upgrade processes.
- Experience or readiness to manage air-gapped, offline deployment environments, including registry and package mirroring.
- A focus on build reproducibility, idempotency, and strict validation procedures to ensure release quality.
- Fluency in written and spoken English communication.
- Note: Deep storage, networking, or GPU technical expertise is not required; the role validates these subsystems but does not architect them.
Preferred Qualifications
- Experience packaging or validating GPU or RDMA-enabled Kubernetes deployments.
- Knowledge of software supply-chain security concepts such as artifact signing, provenance tracking, and Software Bill of Materials (SBOMs).
- Previous involvement in self-hosted or on-premises product distribution workflows.
Success Metrics for First 90 Days
- Automated release-validation gate runs successfully on every release candidate, performing fresh installs and in-place upgrades to ensure readiness.
- Releases are consistently cut from the pipeline controlled by the engineer as a documented, repeatable process.
- The validation gate catches deployment-blocking issues prior to customer delivery.
- Build processes generate reproducible, offline, self-contained artifacts reliably.
- All pipeline operations and release procedures are clearly documented enabling other engineers to manage releases independently.
Level
Senior