Jobgether

AI Research Engineer (Kernel and Inference Optimization)

Jobgether

Remote · Full Time

Be the first to apply

Experience
Any
Salary
—
Openings
1
Posted
1 week ago
Work mode
Work from home
Education
Degree in Computer Science or related field, PhD preferred
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Role Overview

This position involves working at the cross-section of artificial intelligence research, systems engineering, and high-performance model inference, focused on developing and refining model-serving systems across varied hardware platforms. The role offers a balance of hands-on research and low-level engineering with aims to enhance AI system performance, including latency, throughput, memory usage, and scalability, especially on mobile and edge devices.

Responsibilities

  • Architect and implement advanced model-serving frameworks emphasizing high throughput, low latency, and optimized memory use.
  • Create inference pipelines suitable for diverse environments, including mobile and edge systems with restricted resources.
  • Define and strive to meet performance metrics such as response times, token processing rates, throughput, memory consumption, and system dependability.
  • Construct and conduct detailed inference benchmarking in both simulated and production settings to track various performance indicators.
  • Develop and maintain datasets and simulation setups to realistically assess model performance under real-world and constrained conditions.
  • Analyze and resolve computational and memory bottlenecks through techniques like batching, network optimization, and memory management.
  • Implement custom GPU kernels and compute shaders tailored to mobile platforms using Metal Shading Language (MSL).
  • Utilize advanced optimization methods including pruning, quantization, Flash Attention, key-value caching, and speculative decoding approaches.
  • Design and refine distributed inference systems leveraging tensor, pipeline, and expert parallelism for large-scale GPU workloads.
  • Collaborate with diverse engineering and research teams to embed optimized inference solutions into production and edge deployments.
  • Establish evaluation protocols, document tests, benchmark against standards, and iteratively enhance optimization strategies.
  • Monitor live production performance to identify further scalability and efficiency improvements.

Candidate Requirements

  • A degree in Computer Science or related fields; a PhD in NLP, Machine Learning, or a similar specialization with published AI research is highly valued.
  • Strong proficiency in Metal Shading Language, capable of crafting custom compute shaders from the ground up.
  • Experience optimizing low-level kernels and inference workflows on mobile or constrained devices.
  • A proven record of enhancing inference latency, throughput, and memory footprint in specialized domains.
  • Comprehensive knowledge of modern model-serving platforms, inference engines, and advanced deployment optimization methods.
  • Hands-on experience writing GPU kernels for mobile devices such as smartphones.
  • Skill in developing full inference pipelines that progress from model optimization to real-world integration on limited hardware.
  • Ability to leverage empirical research and methodical experimentation to address latency, performance, and memory bottlenecks.
  • Expertise in building evaluation and benchmarking frameworks for inference systems.
  • Familiarity with distributed inference paradigms like tensor parallelism, pipeline parallelism, and expert parallelism suitable for GPU clusters.
  • Deep understanding of diffusion model and Vision Transformer architectures and theoretical underpinnings.
  • Knowledge of cutting-edge inference optimization techniques such as pruning, quantization, Flash Attention, KV Cache optimization, and speculative decoding like EAGLE.
  • Strong analytical problem-solving skills able to dissect complex system limitations and translate research insights into engineering enhancements.
  • Excellent English communication skills with the ability to effectively contribute within remote, cross-disciplinary technical teams.

Benefits and Work Environment

  • Engagement with state-of-the-art AI systems encompassing model serving, inference optimization, mobile and edge computing, and distributed large-scale inference.
  • Remote-first work culture in an international, collaborative team setting.
  • Access to pioneering AI research combined with practical systems engineering challenges.
  • Opportunity to impact critical performance infrastructure measurable in real-world AI use cases.
  • A blend of research-led experimentation alongside hands-on technical development.
  • Involvement with advanced model families including diffusion models, Vision Transformers, and multimodal AI systems.
  • Exposure to complex problems in GPU kernel development, inference engines, memory management, and distributed computing.

Additional Information

This opportunity is presented on behalf of a partner company managing applications and follow-up processes directly. The recruitment uses AI-assisted candidate matching for rapid, unbiased evaluation of applications. Final hiring decisions rest with the partner firm's internal teams.

Applicants' personal data will be processed in compliance with relevant data protection laws, with support for exercising data rights. AI tools may assist in application analysis, although human judgment ultimately guides selection.

Minimum education

Master's Degree

How they work

Teamwork & Collaboration Problem Solving Attention to Detail

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer