Senior DevOps Engineer
London, England, United Kingdom (Hybrid) · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 3 days ago
- Work mode
- Hybrid
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About XYZ Reality
XYZ Reality has pioneered the world's first engineering-grade Augmented Reality solution tailored for the construction sector. Their flagship technology, The Atom, assists construction teams worldwide in achieving more precise, efficient builds with fewer errors.
As a dynamic Series B company expanding operations across the UK, US, and Europe, they require robust infrastructure scaling to support this growth.
Role Overview
We are seeking a Senior DevOps Engineer to take full ownership of the cloud infrastructure and DevOps methodologies that support our BIM Platform. This role entails hands-on engagement with Microsoft Azure, Kubernetes, CI/CD processes, observability tools, security protocols, and developer tooling.
Working closely with Engineering, Data, and Security teams, the candidate will construct and maintain scalable and dependable infrastructure while improving deployment, monitoring, and troubleshooting workflows for developers.
The approach is to have DevOps serve as an enabler rather than a blocker by creating automation, self-service capabilities, and tooling that empower engineering teams to move swiftly and securely without undue dependencies on DevOps resources.
This position is hybrid-based, requiring attendance in the London office three days per week.
Key Responsibilities
- Design, maintain, and advance Microsoft Azure cloud infrastructure alongside Kubernetes clusters.
- Develop and upkeep Infrastructure-as-Code configurations using Terraform, ARM templates, and Helm charts.
- Create and optimize CI/CD pipelines via GitHub Actions to facilitate fast, secure, and repeatable deployments.
- Implement observability solutions encompassing monitoring, logging, alerting, and distributed tracing across infrastructure and services.
- Enhance platform reliability, scalability, and performance through capacity planning, auto-scaling, and resource optimization.
- Establish disaster recovery plans, backups, and failover mechanisms.
- Strengthen security by managing network policies, secrets, and hardening of the platform infrastructure.
- Define and refine incident response workflows, on-call schedules, thorough runbooks, and conduct blameless post-mortem analyses.
- Collaborate with Data teams to support stable and scalable PostgreSQL and MongoDB database systems.
- Build self-service tools, templates, and documentation to reduce unnecessary reliance on DevOps.
- Support infrastructure required for data pipelines and AI/ML model deployment as the technology grows.
- Partner closely with developers to assist in deployment, troubleshooting, and infrastructure best practices.
Required Skills and Experience
- Proven hands-on experience in DevOps, Infrastructure or Cloud Engineering roles.
- Advanced expertise in Microsoft Azure cloud platform.
- In-depth knowledge of Kubernetes including cluster architecture, networking, storage, upgrades, and troubleshooting.
- Extensive experience with Infrastructure-as-Code tools such as Terraform and ARM templates.
- Proficiency in designing and maintaining CI/CD pipelines preferably using GitHub Actions.
- Solid understanding of containerization technology, especially Docker.
- Experience setting up monitoring, logging, alerting, and overall observability within production environments.
- Strong grasp of infrastructure security including network segmentation, encryption, secrets management, and audit logging.
- Capability to build automation and tooling using programming languages like Python, Go, Bash, or TypeScript.
- Excellent troubleshooting capabilities with a willingness to take responsibility for production infrastructure and incident management.
- Ability to make informed, pragmatic technical decisions and clearly communicate infrastructure trade-offs.
Additional beneficial skills include experience with AWS or GCP environments transferable to Azure, high-availability infrastructure design, PostgreSQL and MongoDB management, AI/ML infrastructure, ETL/data pipeline support, FinOps, and compliance with SOC 2 or ISO 27001 standards.
Why Work With Us
- Participate in a Series B company at an exciting international growth phase.
- Hybrid work model based at our London office.
- 25 days of annual leave plus public holidays.
- Private healthcare coverage through Vitality.
- Additional leave during Christmas shutdown.
- Biannual salary reviews.
- Social events during summer and Christmas.
- Complimentary Thursday lunches and post-work gatherings.
- Employee referral incentives.
- Cycle to Work scheme supporting sustainable commuting.
- Opportunity to impact the construction industry through innovative technology.
Additional Information
This role demands a proactive engineer who relishes automating repetitive tasks, taking ownership of infrastructure challenges, questioning inefficient processes, and crafting tools that enhance overall engineering efficiency.
Level
Senior