- Experience
- 10+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 hour ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Role Overview
This opportunity is offered by our partner company seeking a Staff Engineer specializing in Core systems and MLOps, located in Saudi Arabia. The role involves designing foundational infrastructure that supports large-scale web data products and distributed engineering teams. You will be responsible for architecting the core control and context planes that enable reliable operation of services and AI-driven workflows across Kubernetes, Kafka, Java, Python, gRPC, and multi-cloud environments.
Key Responsibilities
- Design and enhance core control and context planes, including service and schema registries, SLO enforcement, and health-based routing, alongside automated canary releases and operational feedback mechanisms.
- Manage the service chassis and golden path by maintaining and upgrading Java and Python multi-language client libraries, standardized workload definitions, Helm charts, and deployment pipelines.
- Define and oversee inter-service contracts such as gRPC and Protocol Buffer specifications, API gateway transcoding, versioning schemes, and schema evolution standards.
- Operate and refine platform infrastructure components including Kubernetes, Terraform, HAProxy/Nginx, Confluent Kafka, real-time billing pipelines, Valkey, and engage in database modernization efforts.
- Lead architectural directions via Requests for Discussion (RFDs) for workflow orchestration, gateway orchestration, multi-cluster routing, automated failover, and critical platform projects.
- Implement and uphold reliability engineering practices like setting SLOs, SLIs, error budgets, fault isolation strategies, and automated weighted canary deployments.
- Participate in on-call rotations for shared infrastructure, lead post-incident reviews, and convert learnings into platform enhancements.
- Mentor engineers across squads, review architectural designs, and promote engineering standards for consistent and reliable software development.
Qualifications & Experience
- Over 10 years of experience in building scalable distributed backend systems with a proven history of establishing internal platforms or libraries widely adopted within engineering teams.
- Expert-level Java skills, including experience with reactive frameworks like Vert.x or Netty, paired with solid Python programming capabilities.
- In-depth knowledge of gRPC and Protocol Buffers with expertise in schema evolution and backward compatibility for critical systems.
- Direct experience managing Kubernetes at scale, Terraform automation, and event-streaming technologies such as Kafka.
- Background in designing automated telemetry systems, materialized views, feature stores, or feedback loops that use live production data to enhance system behavior.
- Strong foundation in reliability engineering encompassing SLO/SLI definition, blast-radius analysis, fault tolerance, and strict service contracts.
- Excellent written communication skills for conveying complex architectural ideas and facilitating team alignment in a distributed remote setting.
- A continuous learner eager to explore new technologies, architectural paradigms, and engineering methodologies.
- Additional advantages include experience with Temporal, DBOS or similar durable execution platforms; MLOps skills such as model serving, performance monitoring, and drift detection; and familiarity with zero-trust networking and service meshes like SPIRE, mTLS, Cilium, Istio, or Envoy.
- Experience developing developer tools such as CLIs, SDKs, or project generators is a plus.
- Background in large-scale web scraping/crawling or contributions to distributed systems and data-extraction open-source projects is desirable.
Benefits
- Fully remote work with flexible hours tailored to your productivity preferences.
- Opportunity to engage in work supporting core infrastructure behind large-scale web data pipelines and distributed systems.
- Access to cutting-edge open-source technologies and evolving AI and web data infrastructure.
- Chance to participate in industry conferences and network with a global, diverse engineering community.
- A culture that values autonomy and trust, with influence over platform architecture and engineering standards across multiple teams.
Additional Information
This position is coordinated by a partner company that manages applications and subsequent recruitment steps. The hiring process employs AI-supported matching to ensure objective candidate evaluation, followed by internal final decisions. Candidate data will be processed respecting applicable data privacy laws, including GDPR. AI tools may assist in application review but do not replace human judgment in hiring decisions.
Level
Mid