- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 2 weeks ago
- Work mode
- In office
- Eligibility
- Applicants must have unrestricted legal work rights in Australia; visa sponsorship is not provided.
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Maincode
Maincode is a pioneering AI research and engineering firm based in Melbourne, dedicated to developing Matilda, Australia's advanced large language model. We operate our own GPUs within local data centers and develop every layer on top, including training and deployment systems, core model programming, assessment methods, and user-facing products. Currently, we are both training the next Matilda iteration and providing access to the current version for public use.
Role Overview
As an AI Researcher, you will collaborate closely with the team to shape the development of the upcoming Matilda model. Your responsibilities will encompass designing architectural and training modifications, conducting scaled experiments on our computational cluster, interpreting results through log analysis, and making informed design decisions. This position requires active coding and engineering involvement.
Key Responsibilities
- Devise and evaluate changes in architecture, objectives, and training procedures for a state-of-the-art large language model.
- Manage large-scale controlled experiments and accurately diagnose cause-and-effect relationships in outcomes.
- Analyze failure points in reasoning, generalization, and representation within models.
- Create evaluation frameworks that focus on capability and reliability rather than mere benchmark performance.
- Compose concise internal reports that transform experimental data, including negative findings, into actionable design choices.
Required Expertise
We seek candidates with a profound understanding of transformer mechanisms, including attention variations, positional encoding techniques, normalization methods, optimization algorithms, tokenization strategies, and data mixture impacts on loss metrics and model behavior. Capability to interpret training logs and familiarity with pre-training and evaluation research is essential, with a demonstrated willingness to challenge and test prevailing theories. Proficiency in production-quality Python programming and frameworks such as PyTorch or JAX is expected. Experience in multi-node distributed training is advantageous but not mandatory.
What This Role Does Not Include
This role excludes product-focused research, prompt engineering, fine-tuning of existing external models, or tasks primarily measured by academic publications.
Additional Information
This is a full-time, in-person position based in Melbourne. Candidates must possess valid and unrestricted work authorization within Australia, as visa sponsorship is not available.