- Experience
- 3+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 2 days ago
- Work mode
- In office
- Education
- Bachelor's degree preferred
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Position
The Cloud Platform Engineer role focuses on designing, building, and enhancing a distributed computing system that supports extensive human and simulation data production workflows.
Key Responsibilities
- Engineer the distributed computing infrastructure enabling human and simulation data processing.
- Create and manage cross-region compute resource scheduling systems, task queues, quota allocations, prioritization, and dynamic scaling mechanisms.
- Decompose sequential workflows into discrete, parallelizable compute tasks that are isolated, retryable, and monitorable.
- Support various execution paradigms including batch, micro-batch, and streaming, while bolstering fault tolerance and resource governance.
- Continuously enhance metrics such as task throughput, queue latency, failure rates, GPU usage efficiency, and compute unit costs.
- Develop clear and observable interfaces for task management accessible to both agents and system components.
Required Qualifications
- Minimum of three years' professional experience in distributed computing, task scheduling, infrastructure, or data platform development.
- Preferably hold a bachelor's degree or higher; however, exceptional candidates may be considered without formal degrees.
- Strong expertise in task scheduling concepts, concurrency, queues, resource pool management, isolation, and fault recovery, with practical production environment experience.
- Background in cloud computing platforms, GPU or compute accelerator technologies, and large-scale asynchronous task orchestration systems.
- Skilled at identifying boundaries of parallelizable units, managing data dependencies, and enforcing resource isolation in complex pipelines.
- Hands-on capability in designing, implementing, debugging, and maintaining production-grade systems.
- Familiarity with AI-enhanced coding and testing tools such as Cursor, Codex, or Claude Code, including the capability to validate AI-generated content.
Preferred Expertise
- Familiarity with technologies and platforms like Ray, Apache Spark, Apache Flink, Kubernetes Jobs, Argo workflows, Slurm schedulers, or GPU task scheduling systems.
- Experience optimizing cross-region scheduling strategies and improving GPU resource utilization in large-scale asynchronous computing environments.
- Background in media processing, point cloud data handling, simulation workloads, or machine learning computational platforms.
Equal Opportunity
The organization is dedicated to fostering an inclusive and diverse workplace culture for all employees.
Minimum education
Bachelor's Degree
Skills
Tools & software
Kubernetes
required