AI Engineer - Inference
Sydney, New South Wales, Australia · Full Time
Be the first to apply
- Experience
- 5+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Firmus Technologies
Firmus Technologies, established in Australia in 2019, is a pioneering company dedicated to developing and delivering sustainable and efficient AI infrastructure across the Asia Pacific region. We engineer and operate innovative AI Factories featuring advanced liquid cooling, energy management, AI orchestration, and construction technologies. Our Firmus AI Cloud platform provides large-scale, energy-conscious GPU cloud computing designed to help developers and organizations efficiently train and deploy AI models at scale while reducing costs.
As a partner of NVIDIA Cloud and Engineering in Asia Pacific, we offer employees the opportunity to contribute to shaping the AI industry with accessible leadership, fast decision-making, and a chance to own impactful projects that integrate closely with energy grid assets benefitting communities.
Role Summary
The Senior AI Engineer focusing on Inference will spearhead the development, operation, and continual enhancement of the AI & Applications team's inference capabilities. This includes establishing a robust engineering base for hosting models internally with self-managed endpoints that are reliable, secure, scalable, and high performing. The role involves onboarding models, provisioning endpoints, managing runtimes, benchmarking performance, and lifecycle management to ensure efficient and predictable access to AI models while controlling cost, security, and infrastructure use.
Key Responsibilities
- Develop and manage self-hosted AI inference services for internal users, customers, and future inference-as-a-service products.
- Design and standardize workflows for model intake, validation, packaging, runtime decisions, deployment, endpoint setup, testing, release, and lifecycle processes.
- Securely provision scalable inference endpoints supporting diverse AI patterns including interactive generation, retrieval-augmented generation, embeddings, reranking, batch processing, multimodal models, tool integration, and agentic workflows.
- Create reusable deployment assets like templates, APIs, SDKs, and configuration protocols to empower users with self-service model endpoint management.
- Leverage leading inference platforms and frameworks such as TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, CUDA, and related tools.
- Optimize inference workloads through techniques like quantization, batching, request routing, caching, load balancing, memory management, and distributed parallelism.
- Produce and validate inference recipes specifying compatible models, runtimes, precision, GPU setups, scaling, and performance benchmarks.
- Apply advanced quantization methods (NVFP4, FP8, INT8), compilation, and optimization while meeting model quality targets.
- Architect distributed inference systems incorporating tensor, pipeline, expert, context, and data parallelism appropriately.
- Collaborate with Kubernetes and scheduling teams to define resource profiles, placement, scaling, quotas, and policies for endpoints.
- Feed workload characteristics and inference metrics into Model-to-Grid products to enhance scheduling and operational decisions.
- Build benchmarking and qualification pipelines encompassing load, latency, concurrency, scaling, regression, and profiling tests.
- Monitor and enhance inference metrics such as latency per token, throughput, GPU and memory utilization, cache hits, scaling, energy use, and cost effectiveness.
- Implement and automate performance regression testing across various system and software upgrades.
- Establish comprehensive observability for inference endpoints including availability, error rates, utilization, capacity, and SLA adherence.
- Support agentic applications by delivering specialized inference endpoints for reasoning, retrieval, tool use, summarization, recommendation, and autonomous workflows.
- Expose inference performance and capacity data to agentic systems to facilitate intelligent decision-making and optimization.
- Work cross-functionally with Product, UX, DevOps, Security, Infrastructure, and Operations to ensure clarity, security, and operational robustness across inference services.
Skills and Experience
- Over 5 years of software engineering experience, with at least 3 years specifically in AI inference, model serving, ML systems, or comparable performance-critical computing environments.
- Proven track record building or significantly improving production model-serving platforms, inference APIs, GPU-accelerated services, or multi-tenant AI developer platforms.
- Hands-on expertise with modern inference frameworks like TensorRT-LLM, Triton, NVIDIA Dynamo, or Hugging Face Text Generation Inference.
- Deep understanding of NVIDIA AI stack: CUDA, cuDNN, NCCL, TensorRT, GPU profiling, and distributed GPU communication.
- Comprehensive knowledge of LLM inference behavior including prompt handling, token generation, batching, concurrency, KV cache, scheduling, routing, and latency/throughput optimization.
- Experience with model optimization techniques such as quantization, calibration, mixed precision, kernel fusion, memory and cache optimizations, speculative decoding, parallelism, and accuracy validation.
- Strong proficiency in Python plus working knowledge of C++ or Go applied to inference services, automation, benchmarking, and performance-critical software development.
- Familiarity with distributed inference/training strategies including tensor, pipeline, expert, context, data parallelism, collective communications, fault tolerance, and scaling across nodes.
- Operational knowledge of Kubernetes, containers, CI/CD pipelines, GitOps, autoscaling, workload scheduling, multi-tenant environments, and observability tools.
- Understanding of GPU infrastructure specifics like topology, NVLink, PCIe, NUMA, network affinity, RDMA, RoCEv2, network fabrics, and storage bandwidth affecting inference efficiency.
- Experience creating benchmarking tests for inference performance including throughput, latency, resource utilization, scaling efficiency, power consumption, and cost analysis.
- Awareness of inference use cases such as retrieval augmented generation, embeddings, reranking, multimodal inference, model routing, tool calling, and agentic applications.
- Knowledge of security and governance for inference services including authentication, authorization, tenant isolation, quota enforcement, secrets management, audit logging, abuse prevention, and data privacy.
Success Metrics
- Availability of reliable, secure, scalable inference endpoint services supporting internal priorities, customer products, and Model-to-Grid integrations.
- Significant reduction in model onboarding, validation, optimization, deployment, and lifecycle management time.
- High adoption of standardized workflows, catalogues, templates, APIs, SDKs, recipes, and self-service capabilities for inference deployment.
- Strong foundational systems enabling future inference-as-a-service features including multi-tenancy, quota controls, usage monitoring, observability, support, and release governance.
- Performance improvements across inference metrics such as latency, throughput, GPU and memory efficiency, and scaling behaviors.
- Lowered operational cost parameters including cost per request and energy use while maintaining quality and uptime standards.