J

AI Engineer, Inference

Jobverse.io

Launceston, Tasmania, Australia · Full Time

Be the first to apply

Experience
Any
Salary
—
Openings
1
Posted
1 week ago
Work mode
In office
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Overview

Firmus Technologies, a prominent AI infrastructure provider across Asia Pacific, is searching for a Senior AI Engineer specializing in inferencing. This role involves creating and enhancing self-hosted AI inference systems, encompassing tasks from model onboarding through endpoint provisioning to runtime optimization, benchmarking, monitoring, security, and life cycle management. The position also contributes to Model-to-Grid products and agentic software by delivering dependable, scalable, and performance-focused model endpoints.

Key Responsibilities

  • Develop and maintain self-hosted AI inference services for internal uses, client products, and upcoming Inference-as-a-service offerings.
  • Create and enforce standardized model onboarding workflows, covering aspects such as model intake, compatibility checks, packaging, runtime choice, optimization, deployment, registration, testing, releases, and lifecycle oversight.
  • Set up and oversee secure, scalable inference endpoints tailored for interactive content generation, retrieval-augmented generation (RAG), embeddings, reranking, batch analysis, multimodal tasks, tool integrations, and agent-like workflows.
  • Design reusable deployment blueprints, APIs, SDKs, configuration standards, and enable users with self-service tools for managing model endpoints.
  • Work extensively with leading inference frameworks and toolkits, including TensorRT-LLM, TensorRT, SGLang, vLLM, Triton Inference Server, and NVIDIA technologies like Dynamo, NIM, CUDA, cuDNN, NCCL, plus performance and profiling instruments.
  • Enhance model-serving efficiency through strategies like quantization, compilation, batch processing, request routing, KV-cache and prefix caching, speculative decoding, load balancing, route optimization, memory management, and multi-node parallelism.
  • Formulate and validate inference recipes outlining compatible model and framework versions, precision and GPU configurations, topology prerequisites, scaling protocols, scheduler profiles, benchmark outcomes, and expected performance standards.
  • Utilize optimization techniques including NVFP4, FP8, INT8 quantization, TensorRT compilation, kernel tuning, effective attention mechanisms, and memory strategies without compromising model quality targets.
  • Design distributed inference setups for large models leveraging tensor, pipeline, expert, context, and data parallelism methodologies as required.
  • Collaborate with Kubernetes and scheduler teams to specify endpoint resource demands, placement rules, topology preferences, priority classes, quotas, autoscaling, capacity allocation, and workload governance policies.
  • Provide inference workload data, benchmarks, and performance insights for Model-to-Grid products to enhance endpoint scheduling, placement, capacity planning, and operational decisions.
  • Build benchmarking and validation procedures employing controlled tests, reproducible baselines, load and latency tests, throughput, concurrency, scaling tests, profiling, regression testing, and standard benchmarks.
  • Track and improve critical inference metrics such as time-to-first-token, inter-token latency, tokens and requests per second, end-to-end latency, concurrency, GPU and memory usage, cache efficiency, scaling, power consumption, and cost-effectiveness.
  • Implement automated regression tests and release qualification for all model versions, runtime/toolkit updates, CUDA/driver changes, Kubernetes or scheduler modifications, network/storage upgrades, and new GPU platforms.
  • Establish observability tools for inference services monitoring endpoint availability, request volumes, latencies, queuing, error rates, GPU metrics, cache performance, capacity, costs, power use, and service-level objectives.
  • Partner with the agentic apps team to provide dedicated self-hosted endpoints optimized for tasks such as agent planning, retrieval, tool invocation, summarization, diagnostics, recommendations, optimizations, and AI-factory functions.
  • Deliver governed information on inference, benchmarking, recipes, performance, and capacity to agentic systems for model selection, degradation detection, bottleneck identification, optimization planning, and validation.
  • Coordinate cross-functionally with Product, UX, DevOps, Platform, Infrastructure, Security, and Global Operations to ensure inference provisioning, selection, configuration, monitoring, quota governance, and troubleshooting are transparent, secure, and maintainable.
🤖
Online · instant AI help
Broxer