vCluster

AI Infrastructure Engineer

vCluster

Remote · Full Time

Be the first to apply

Experience
5+ yrs
Salary
USD 140,000 – USD 165,000 / year
Openings
1
Posted
2 weeks ago
Work mode
Work from home
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

Join vCluster as an AI Infrastructure Engineer, directly partnering with customers in the pivotal early phases of deploying GPU-accelerated Kubernetes environments. This position extends beyond typical professional services by engaging in pre-sales proof of value projects designed to progress clients from bare metal GPU nodes to fully operational clusters. You will pioneer technical playbooks and facilitate scalable deployment practices crucial for supporting rapid growth in GPU AI Clouds and AI Factory ecosystems.

Primary Responsibilities

  • Lead comprehensive technical rollouts for GPU-focused neocloud and AI Factory customers, starting from bare metal setups to validated vCluster environments.
  • Optimize infrastructure by configuring and troubleshooting bare metal GPU nodes, including setting up CNI networking, GPU Operators, storage platforms, and RDMA/InfiniBand.
  • Deploy and verify Kubernetes and vCluster solutions to enable GPU-powered managed Kubernetes services.
  • Enable customer's technical teams to efficiently operate and expand their platforms independently through knowledge transfer.
  • Create detailed, reusable playbooks and documentation to facilitate streamlined deployment processes for future clients.
  • Collaborate closely with Engineering and Product groups to identify recurring challenges and influence the product roadmap based on field feedback.
  • Support Sales during pre-sale phases requiring deep infrastructure expertise to establish meaningful proofs of value.

Required Qualifications & Skills

  • Extensive production experience (5+ years) managing Kubernetes, particularly on bare metal or other highly complex infrastructures.
  • Proficient with NVIDIA GPU Operator, CUDA tooling, and systems-level GPU node configuration.
  • Strong grasp of networking including CNI plugins, overlay networks, load balancing, and diagnosing complex connectivity issues.
  • Experience with persistent volume management, CSI drivers, and distributed storage systems such as Ceph, Rook, Weka, or Longhorn.
  • Adept at navigating fast-changing, ambiguous environments while independently creating operational documentation.
  • Embrace modern technology stacks over legacy systems and handle diverse technical challenges across internal services and deployment pipelines.

Preferred Additional Skills

  • Automation scripting using Bash, Python, or Go.
  • Kubernetes certifications such as CKA or experience developing Kubernetes Operators.
  • Acquaintance with AI/ML workloads, including inference serving and GPU scheduling for large language models.
  • Experience contributing to AI automation documentation for shared knowledge bases.

About vCluster Labs

vCluster Labs is an innovative, venture-backed startup pioneering Kubernetes virtualization technology tailored for AI workloads. With over $30 million in funding from premier investors and a remote-first global team, we empower AI cloud providers and enterprises to manage GPU infrastructure effortlessly at scale. Our open-source project, vCluster, supports tens of millions of virtual clusters and has become a cornerstone in Kubernetes-native virtualization for AI applications.

Benefits

  • Competitive salary with equity options.
  • Comprehensive insurance options including health, dental, vision, and life coverage for employees and dependents, varying by location.
  • Flexible working hours accommodating personal schedules without rigid clock-in requirements.
  • Flexible work location policies to accommodate life changes and remote preferences.

Culture and Values

  • Make it Happen: Demonstrate relentless determination and prioritize impactful actions to overcome challenges.
  • Own the Outcome: Take full responsibility for delivering value beyond completing tasks.
  • Create Wow: Strive for exceptional experiences internally and externally by supporting colleagues and customers.
  • Open Source, Open Mind: Foster meritocracy and contribute actively to open-source projects, valuing ideas regardless of origin.
  • Build Tomorrow’s Standards, Intentionally: Innovate boldly while respecting the discipline required for mission-critical infrastructure.

Compensation

The salary range for this role is from $140,000 to $165,000 per year.

How they work

Teamwork & Collaboration Problem Solving Adaptability Independence
🤖
Online · instant AI help
Broxer