Jobgether

AI Research Engineer (Kernel & Inference Optimization)

Jobgether

Remote · Full Time

Be the first to apply

Experience
Any
Salary
—
Openings
1
Posted
2 hours ago
Work mode
Work from home
Education
Bachelor's degree in Computer Science or related field; PhD preferred
Resume
Required to apply

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Job Overview

We are seeking an AI Research Engineer specializing in kernel development and inference optimization to join a partner company based in the United Arab Emirates. This role lies at the crossroads of AI research, systems engineering, and efficient model inference across various hardware profiles including mobile and edge devices.

Primary Responsibilities

  • Design and implement high-performance model-serving architectures optimized for low latency, high throughput, and efficient memory usage.
  • Develop and maintain inference pipelines suited for diverse platforms, including resource-limited mobile and edge hardware.
  • Set and monitor performance metrics for latency, token generation speed, throughput, memory consumption, and system reliability.
  • Construct and run thorough inference benchmarks in both simulated and live environments, tracking errors and resource utilization.
  • Create benchmark datasets and simulation environments that reflect practical and constrained use cases.
  • Identify and resolve computational and memory bottlenecks by enhancing batching, networking, memory management, and system-level performance.
  • Author custom GPU kernels and compute shaders targeting mobile hardware, specifically using Metal Shading Language (MSL).
  • Apply leading inference optimization strategies such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.
  • Design and refine distributed inference systems utilizing tensor parallelism, pipeline parallelism, and expert parallelism for large GPU clusters.
  • Collaborate with cross-functional teams to integrate optimized inference frameworks into production and edge applications.
  • Define methods for performance evaluation, document experiments, compare against established standards, and iteratively enhance optimization methods.
  • Continuously monitor deployed systems to identify opportunities for scalability, efficiency, and reliability improvements.

Qualifications

  • Bachelor's degree in Computer Science or a related field; a PhD in NLP, Machine Learning, or allied disciplines with a strong AI research profile is highly desirable.
  • Extensive experience with Metal Shading Language (MSL), including developing custom compute shaders from the ground up.
  • Proven skills in low-level kernel and inference optimization targeting mobile or resource-constrained environments.
  • Track record of significantly improving inference latency, throughput, and memory footprint on domain-specific tasks.
  • In-depth knowledge of modern model-serving architectures and optimization for AI deployment at scale.
  • Hands-on experience writing GPU kernels for mobile platforms like smartphones.
  • Expertise in building production-ready end-to-end inference pipelines on constrained hardware, from model tuning to integration.
  • Strong analytical and experimental approach to solving latency, computation, and memory challenges.
  • Design and execution of rigorous inference system evaluation and benchmarking frameworks.
  • Knowledge of distributed inference technologies, including tensor, pipeline, and expert parallelisms on GPU clusters.
  • Strong familiarity with diffusion model mathematics and Vision Transformer architectures.
  • Hands-on experience with inference optimizations including pruning, quantization, Flash Attention, KV Cache optimization, and speculative decoding like EAGLE.
  • Excellent problem-solving capabilities and the aptitude to translate research into practical engineering solutions.
  • Fluent English communication skills and the ability to work effectively within globally distributed, cross-disciplinary teams.

Benefits

  • Work on cutting-edge AI technologies involving model serving, inference optimization, mobile and edge computing, and large-scale distributed AI inference.
  • Join an international team in a flexible remote-first work environment.
  • Engage with state-of-the-art AI research challenges paired with practical engineering development.
  • Contribute to critical AI infrastructure where performance improvements have tangible real-world impact.
  • Collaborate in a research-driven yet hands-on engineering culture.
  • Experience working with advanced architectures such as diffusion models, Vision Transformers, and multimodal AI systems.
  • Tackle complex technical challenges in GPU kernel development, inference pipelines, memory optimization, and distributed computing.

Additional Information

The application process is managed by a partner organization, which will coordinate all subsequent steps after submission. The role offers an opportunity to engage deeply in AI research while contributing to real-world deployments across diverse computing platforms.

Data privacy and application processing utilize AI-enhanced tools to ensure fair, objective evaluation, while final decisions are made by human teams. Applicants' data is handled according to applicable data protection legislation.

Minimum education

Bachelor's Degree

How they work

Communication Teamwork & Collaboration Problem Solving Attention to Detail

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer