Canva

Senior Machine Learning Engineer - Research Optimisation

Canva

Sydney, New South Wales, Australia (Hybrid) · Full Time

Be the first to apply

Experience
5+ yrs
Salary
Openings
1
Posted
29 నిమిషాలు క్రితం
Work mode
Hybrid
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Company and Location

Join Canva's innovative team focused on transforming the design experience worldwide. This position is based in Sydney, New South Wales, where the office doubles as a vibrant community hub encouraging collaboration and connection. The role offers a hybrid working arrangement, blending remote flexibility with in-person teamwork at the Sydney campus.

Role Overview

As a Senior Machine Learning Engineer specializing in research optimisation, you will serve as the crucial link between the research team and production deployment. Your responsibilities will include converting experimental ML models into stable, scalable user-facing features while enhancing training and inference efficiency in terms of speed, cost, and reliability.

Primary Responsibilities

  • Transform experimental models into production-ready services by refactoring, containerizing, testing, and integrating them within a centralized monorepo for scalable use.
  • Analyze and tune PyTorch-based training workloads to enhance GPU usage and remove bottlenecks related to compute, memory, I/O, and networking.
  • Improve distributed training configurations across multi-GPU and multi-node setups, selecting optimal parallelism strategies tailored to different workloads.
  • Develop and maintain inference services, SDKs, and common libraries standardizing pre-processing and post-processing across model variants.
  • Implement robust rollout frameworks including feature flags, canary deployments, A/B tests, and automated rollback mechanisms for multi-variant release management.
  • Establish continuous integration and deployment pipelines, artifact management, and reproducibility standards for ML models and services.
  • Integrate observability solutions such as metrics, logging, and tracing, establishing SLIs and SLOs to ensure reliability and performance of training and inference workloads.
  • Enhance inference efficiency through batching, caching, quantization, compilation, and hardware optimization to lower latency and operational costs.
  • Collaborate across the training infrastructure including Kubernetes orchestration and high-performance storage systems to streamline research workflows.
  • Partner with researchers and product engineers for code reviews, pair programming, and documentation to accelerate adoption of best practices and libraries.
  • Promote strong engineering practices in the research codebase including testing strategies, dependency management, and eliminating redundant code.

Ideal Candidate Profile

  • Solid foundation in software engineering with proficiency in Python, capable of converting prototypes into production-grade machine learning services.
  • Proven experience deploying ML systems in production environments using containers, APIs, and CI/CD within monorepo setups.
  • Hands-on expertise in optimizing PyTorch training and inference, with skills in profiling GPU workloads and managing compute resources.
  • Comfortable operating in containerized environments and familiar with Kubernetes for debugging and improving ML workloads.
  • Adept at reading and restructuring research code into clean, stable, and well-documented software components.
  • Understanding of service observability and reliability concepts, applying metrics, logging, and tracing to machine learning systems.
  • Holistic understanding of the full technology stack from storage and networking to model code, facilitating communication across cross-functional teams.
  • Strong communicator with an empathetic approach to mentoring and guiding teammates on adopting libraries and best practices.
  • General cloud platform experience, preferably AWS.

Preferred Qualifications

  • Experience with model serving and optimization tools like ONNX, TorchScript, Triton, and quantization techniques.
  • Skills in CUDA kernel programming or use of compilation frameworks such as torch.compile, Triton, and TensorRT to accelerate model performance.
  • Familiarity with distributed training frameworks including FSDP, DDP, DeepSpeed, or Megatron at scale.
  • Knowledge of high-performance storage solutions like Weka, Vast, or Lustre, and understanding their data loading implications.
  • Experience with experimentation platforms supporting feature flags, A/B testing, and safe deployment strategies.
  • Background with multimodal or image generation models and large language model adjacent tooling, advantageous but not core.
  • Awareness of MLOps concepts such as model registries, artifact repositories, and version/dependency management.

Impact and Contributions

You will significantly shorten the journey from research prototype to reliable, user-facing features. By consolidating duplicated efforts, unifying interfaces, enabling parallel experimentation, and enhancing GPU efficiency, your work will empower the Canva platform to scale ML innovations rapidly and cost-effectively.

Additional Information and Benefits

  • Equity participation opportunities, making you a stakeholder in company success.
  • Inclusive parental leave policies supportive of all parents and carers.
  • Annual Vibe & Thrive allowance for wellbeing, social activities, and office setup support.
  • Flexible leave options designed to support personal recharge and social contribution.

Recruitment Process

The hiring process evaluates your experience, skills, and passion along with cultural compatibility. Interviews are conducted virtually and include interactive and real-time challenges that mirror the role’s responsibilities. Certain interviews may incorporate problem-solving using AI tools to examine your approach to technology-augmented problem solving.

Level

Senior

Tools & software

PyTorch required

How they work

Communication Teamwork & Collaboration Problem Solving Leadership

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer