- Experience
- 4+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Fuse Energy
Fuse Energy is a pioneering energy startup dedicated to making energy more abundant and affordable by leveraging first-principles thinking combined with innovative technologies. With over $200 million raised from prominent investors and strategic backers, Fuse is building a fully integrated energy company. Their operations include developing solar energy projects, battery systems, grid infrastructure enhancements, real-time energy trading, AI integration, and distributed energy installations directly to consumers, eliminating intermediaries to lower costs.
Role Overview
In response to the rapid growth in data center energy demand, Fuse Energy is expanding into high-performance computing infrastructure at the nexus of energy and artificial intelligence. The CUDA Engineer will be responsible for writing and refining low-level GPU code that drives inference workloads. This involves crafting custom CUDA kernels, optimizing performance across memory and compute bottlenecks, and maximizing throughput at detailed GPU hardware levels including Streaming Multiprocessors (SMs) and warps.
Key Responsibilities
- Develop and improve customized CUDA kernels specifically targeting core transformer inference functions.
- Conduct profiling to detect and resolve performance issues related to occupancy, memory bandwidth, and warp divergence.
- Apply kernel fusion techniques to minimize memory transfers and overhead in the inference pipeline.
- Optimize memory access strategies and manage memory hierarchy effectively to leverage maximum bandwidth.
- Implement quantization-aware and mixed-precision kernels to enhance latency and decrease memory usage.
- Create and fine-tune caching systems for efficient autoregressive decoding processes.
- Configure kernel launches optimized for specific GPU architectures.
- Benchmark and compare kernels versus existing standards to achieve noticeable improvements in throughput and latency.
- Develop tests to safeguard against performance regressions and maintain code correctness.
- Manage and update internal CUDA libraries, contributing to coding standards and team documentation.
Experience and Skills Required
- Minimum of four years developing production-level CUDA code with proven success in delivering performance-critical kernels.
- Comprehensive knowledge of GPU microarchitecture including warps, occupancy, register pressure, and memory hierarchies.
- Expertise in CUDA C++, including use of streams and asynchronous execution methods.
- Proficient in profiling tools to discern between compute-bound and memory-bound performance bottlenecks.
- Experience in kernel fusion, memory coalescing, and avoiding warp divergence.
- Skill in designing and implementing quantized and mixed-precision kernels.
- Strong understanding of parallel algorithm development and the tradeoffs in numerical precision.
- Additional advantages include experience with transformer or attention mechanisms, autoregressive decoding, building high-performance GPU libraries, HPC or latency-sensitive engineering, multi-GPU or multi-node kernel optimization, and familiarity with PTX or SASS assembly to assess kernel efficiency.
Benefits
- Competitive remuneration package with equity eligibility.
- Biannual bonus program.
- Fully funded technology setup tailored to individual requirements.
- Private health insurance coverage.
- Breakfast and dinner allowances for employees working onsite.
- Note: Benefits may vary depending on global location due to hiring geography.
Industry
Energy