Nebius

Senior Software Engineer - Data Platform (C++)

Nebius

London Area, United Kingdom · Full Time

Be the first to apply

Experience
5+ yrs
Salary
—
Openings
1
Posted
2 hours ago
Work mode
In office
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About Nebius

Nebius is pioneering the evolution of cloud infrastructure tailored for the worldwide AI market. Our full-stack AI cloud platform seamlessly supports developers and enterprises from data management and model training to deployment, eliminating the typical expense and complexity associated with large-scale internal AI/ML infrastructure. Founded by engineers with extensive expertise, our platform tackles challenging problems across computing, storage, networking, and applied AI. Listed on Nasdaq (NBIS) and headquartered in Amsterdam, Nebius operates globally with R&D centers across Europe, the UK, North America, and Israel. Our team exceeds 1,500 members, including hundreds of engineers specializing in hardware, software, and AI research.

Role Overview

We seek a Software Engineer skilled in C++ to contribute to the Nebius Data Platform. This platform serves as a distributed storage and processing system, acting as the definitive data source and backbone for numerous internal and external applications. The platform consolidates storage, compute, and analytics capabilities into a single multi-tenant ecosystem based on the open-source YTsaurus project, enabling us to develop and deploy features internally without upstream delays.

Platform Components

  • Distributed Storage (Cypress): featuring transactional semantics, tiered storage, erasure coding, replication, and high reliability.
  • Compute & ETL: cluster-wide job scheduling across tens of thousands of CPU cores, MapReduce, SQL-like YQL processing, and Spark-over-YTsaurus (SPYT) for modern data workflows.
  • Interactive Analytics (CHYT): On-node ClickHouse® instances delivering low-latency SQL queries on data in place.
  • Dynamic Tables: low-latency NoSQL key-value stores with distributed ACID transactions supporting OLTP and feature store workloads.
  • Orchestracto: Native workflow orchestration similar to Airflow, tightly integrated with the platform.

Key Responsibilities

  • Develop and maintain new features within YTsaurus core using C++, focused on production-level stability and reliability.
  • Enhance platform architecture and operational models to support multi-cluster scaling, shared components, and consistent user experiences across teams and use cases.
  • Optimize the overall platform user experience including APIs, safety mechanisms, debugging tools, and automation workflows for both internal and external stakeholders.
  • Ensure production reliability through incident management including on-call rotations, root cause analysis, and implementation of permanent solutions.

Example Projects

  • Implement sharded YTsaurus masters with Kubernetes operator support to eliminate bottlenecks and enable extensive cluster scaling.
  • Accelerate and stabilize CHYT interactive SQL performance under high loads with techniques such as data skipping and enhanced execution monitoring.
  • Develop Orchestracto into a robust platform product with defined building blocks, developer experience improvements, and governance for workflow sharing.
  • Scale and improve native YTsaurus Parquet-on-S3 workloads by addressing replication, lifecycle consistency, and metadata performance.
  • Design comprehensive audit trails for data modifications across diverse storage and compute systems.

Technical Stack

  • Core development with modern C++20, leveraging asynchronous and multithreading primitives.
  • Supporting services and tools built in Go and Python, including microservices, utilities, and integration tests.

Qualifications

  • Minimum five years of professional software engineering experience.
  • Expert proficiency in C++ programming for core system development.
  • Familiarity with Python and/or Go sufficient to navigate and contribute to related services.
  • Experience building or operating highly concurrent and distributed service systems.
  • Strong production mindset, including competency in remote system access (SSH), log and metrics analysis, and distributed system troubleshooting.
  • Firm grounding in computer science fundamentals including algorithms, data structures, and concurrency principles.

Preferred Experience

  • Background with Big Data technologies such as YTsaurus, Hadoop, Spark, ClickHouse, or Kafka ecosystems.
  • Experience with multi-tenant platform architectures, scheduling, resource management, and reliability engineering.
  • Skills in performance optimization including profiling, reducing lock contention, and balancing latency and throughput.

Interview Process

Coding interviews are part of the recruitment procedure.

Benefits and Culture

  • Competitive salary package.
  • Opportunities for career advancement and professional development.
  • Autonomy and flexible work environment.
  • Collaborative and innovative company culture.
  • Chance to contribute to impactful AI technologies.
  • Diverse international team of talented professionals.

Work Environment Values

At Nebius, the work atmosphere is characterized by rapid progress, bold innovation, continuous growth, meaningful impact, genuine ownership, and the opportunity to influence the future of AI technologies.

Equal Opportunity Employment

Nebius is committed to equal employment opportunities and fostering an inclusive, diverse workplace. We prohibit discrimination based on a wide range of legally protected characteristics. All applicants must be authorized to work in their respective country and provide proof of eligibility if hired. Accommodations are available during the application process upon request.

How they work

Teamwork & Collaboration Problem Solving Attention to Detail

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer