L

Cloud Platform Engineer

Lightwheel

Singapore · Full Time

Be the first to apply

Experience
3+ yrs
Salary
Openings
1
Posted
2 days ago
Work mode
In office
Education
Bachelor's degree preferred
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Position

The Cloud Platform Engineer role focuses on designing, building, and enhancing a distributed computing system that supports extensive human and simulation data production workflows.

Key Responsibilities

  • Engineer the distributed computing infrastructure enabling human and simulation data processing.
  • Create and manage cross-region compute resource scheduling systems, task queues, quota allocations, prioritization, and dynamic scaling mechanisms.
  • Decompose sequential workflows into discrete, parallelizable compute tasks that are isolated, retryable, and monitorable.
  • Support various execution paradigms including batch, micro-batch, and streaming, while bolstering fault tolerance and resource governance.
  • Continuously enhance metrics such as task throughput, queue latency, failure rates, GPU usage efficiency, and compute unit costs.
  • Develop clear and observable interfaces for task management accessible to both agents and system components.

Required Qualifications

  • Minimum of three years' professional experience in distributed computing, task scheduling, infrastructure, or data platform development.
  • Preferably hold a bachelor's degree or higher; however, exceptional candidates may be considered without formal degrees.
  • Strong expertise in task scheduling concepts, concurrency, queues, resource pool management, isolation, and fault recovery, with practical production environment experience.
  • Background in cloud computing platforms, GPU or compute accelerator technologies, and large-scale asynchronous task orchestration systems.
  • Skilled at identifying boundaries of parallelizable units, managing data dependencies, and enforcing resource isolation in complex pipelines.
  • Hands-on capability in designing, implementing, debugging, and maintaining production-grade systems.
  • Familiarity with AI-enhanced coding and testing tools such as Cursor, Codex, or Claude Code, including the capability to validate AI-generated content.

Preferred Expertise

  • Familiarity with technologies and platforms like Ray, Apache Spark, Apache Flink, Kubernetes Jobs, Argo workflows, Slurm schedulers, or GPU task scheduling systems.
  • Experience optimizing cross-region scheduling strategies and improving GPU resource utilization in large-scale asynchronous computing environments.
  • Background in media processing, point cloud data handling, simulation workloads, or machine learning computational platforms.

Equal Opportunity

The organization is dedicated to fostering an inclusive and diverse workplace culture for all employees.

Minimum education

Bachelor's Degree

Tools & software

Kubernetes required

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer