T

GPU Performance Engineer (Inference)

TensorX

Remote · Full Time

Be the first to apply

Experience
Any
Salary
—
Openings
1
Posted
1 hour ago
Work mode
Work from home
Education
BSc/MSc/PhD in Computer Science or related technical discipline or equivalent practical experience
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About TensorX

TensorX is a Dublin-based sovereign AI infrastructure platform running advanced open-weight large language models on proprietary NVIDIA Blackwell GPUs across European datacentres, operating under EU jurisdiction. Their OpenAI-compatible API ensures customer data privacy with no retention after request completion. They serve regulated industries including finance, healthcare, government, developers, and AI platforms, prioritizing AI adoption while maintaining compliance, data privacy, and high performance.

Role Overview

As a GPU Performance Engineer on the inference team reporting directly to the CTO, you will optimize GPU inference engines to maximize throughput within latency targets. You will independently manage and resolve engine issues, patch and enhance open-source inference engines, and delve into GPU kernel-level tuning when engine bottlenecks occur. Your contributions will directly affect how many requests each GPU can serve under real production conditions rather than synthetic benchmarks.

Responsibilities

  • Manage the full lifecycle of engine issues—from identification and replication through debugging, patching, and deployment—while documenting findings thoroughly.
  • Diagnose complex performance problems that cross GPU memory, engine scheduler, containers, router, and traffic patterns, providing well-informed analyses rather than mere queries.
  • Identify and recover lost throughput (goodput) per GPU across kernel executions, scheduler behavior, routing, or configurations using production workloads for evaluation.
  • Patch and maintain open-source inference engines like SGLang and vLLM; contribute fixes upstream whenever possible.
  • Develop or improve GPU kernels related to Blackwell GPU attention mechanisms, FP8/FP4 precision paths, and memory-bound decoding, based on profiling insights.
  • Lead new model integrations (day zero bring-up), conducting A/B tests on parallelism layouts, KV cache, and decoding parameters to optimize configurations for inference team usage and resource planning.
  • Optimize caching mechanisms, including prefix caching, KV cache operations, and replication strategies under tensor parallelism, collaborating on cache-aware routing enhancements.
  • Co-own the pre-production validation gate alongside the Inference Team to ensure all changes are rigorously assessed before reaching customers.
  • Produce comprehensive technical documentation recording performance investigations, trade-offs, and benchmarking methodologies accessible to both technical staff and customers.

Required Skills and Experience

  • Demonstrated ability to analyze inference engine source code to uncover root causes of issues, ideally supported by contributions such as patches, reports, or post-mortems.
  • Strong grasp of GPU architecture and CUDA concepts, including memory hierarchy, occupancy, and kernel execution characteristics (compute-bound vs. memory-bound).
  • Familiarity with transformer inference mechanisms, such as various attention types (MLA, DSA, GQA), KV cache behaviors, batching strategies, and quantization impact on accuracy.
  • Proficiency in Python programming, plus the ability to read and understand C++ and CUDA codebases.
  • A methodical, measurement-driven approach to testing changes, with clear experimental design and rigor in correctness verification before trusting performance data.
  • Commitment to transparency about results, valuing negative outcomes as important findings.
  • Keen enthusiasm for continuous learning in a rapidly evolving technology stack, prioritizing source code familiarity over previous knowledge.
  • Comfort using AI-powered coding tools such as Claude Code and Codex to accelerate development.
  • Excellent communication skills, capable of conveying technical complexities clearly to both specialized and non-technical audiences in ambiguous environments.

Preferred Qualifications

  • Experience writing and benchmarking CUDA kernels targeting Hopper or Blackwell GPUs.
  • Prior contributions to open-source projects like SGLang, vLLM, or TensorRT-LLM.
  • Familiarity with GPU kernel development tools including Triton, CUTLASS, CuTe, or ThunderKittens.
  • Recognition in GPU kernel performance competitions or challenges such as GPU Mode leaderboards, MLSys contests, or FlashInfer.
  • Hands-on experience profiling large model memory usage and resolving out-of-memory errors.
  • Practical knowledge of advanced Blackwell GPU features like tcgen05, TMEM, and TMA.
  • Experience creating public technical writing.

Why Join TensorX?

  • Work directly on the inference engine and GPU kernel level, influencing the core margin-driving components not typically accessible in other environments.
  • Operate on TensorX's own dedicated NVIDIA B300 GPUs across Dublin and Helsinki, not shared cloud resources.
  • Engage with high-concurrency, long context production workloads that pose unique challenges beyond standard benchmarking.
  • See measurable impact with better fleet utilization, achieving approximately double the requests handled within latency targets compared to standard configurations.
  • Contribute to open-source engines with the chance for co-authorship in research publications.
  • Explore a broad research agenda covering KV cache beyond GPU memory, context parallelism, expert parallelism scales, and attention/feed-forward disaggregation.

Education & Qualifications

BSc, MSc, or PhD in Computer Science, Engineering, Machine Learning, or related technical fields, or equivalent proven skills and accomplishments.

Additional Information

  • Highly competitive compensation reflecting experience.
  • 25 days of paid annual leave.
  • Hybrid work model based from central Dublin office with flexibility for remote work.
  • Free inference tokens provided.
  • No agency assistance required.

Minimum education

Doctorate

Tools & software

How they work

Communication Problem Solving Learning Agility

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer