- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Eligibility
- Applicants must have permanent work rights in Australia as visa sponsorship is not provided.
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
Matilda is Australia's large language model (LLM), whose performance depends heavily on the quality of its training data. The Signal Engineer plays a crucial part in determining the quality ceiling of the model by creating and maintaining pipelines that convert massive, raw, and messy datasets into well-prepared training material. This position blends engineering expertise with editorial decision-making applied through programming.
Key Responsibilities
- Develop and maintain scalable data pipelines that ingest, cleanse, deduplicate, filter, and score training datasets ranging from terabytes to petabytes.
- Create quality assessment tools, classifiers, and heuristics to isolate valuable data from irrelevant information.
- Design dataset mixtures and conduct experiments to evaluate their impact on model improvement.
- Build tools for exploring, sampling, and auditing the contents of the data corpus.
- Collaborate closely with research and training engineering teams to ensure data selection aligns with desired model behaviors.
Candidate Requirements
- Proven engineering skills with experience in Python, data tooling, distributed processing, and creating reliable pipelines.
- Exceptional attention to detail given the compounding effect of minor errors at large data scales.
- Strong judgment and discernment regarding what constitutes high-quality training data.
- Experience handling extremely large and unstructured datasets.
- A keen curiosity about the influence of data on model functionality.
- High capacity for rapid learning; advanced degrees or prior LLM experience are not mandatory.
Preferred Qualifications
- Practical exposure to working with web-scale datasets or pretraining data workflows.
- Experience processing unstructured text data.
- Familiarity with distributed data processing frameworks such as Spark or Ray.
- Knowledge of data deduplication, quality classification, and tokenization techniques.
Additional Information
This is a full-time position based physically in Melbourne, Australia, collaborating closely with our onsite team. Visa sponsorship is not available; applicants must have valid and unrestricted work authorization for Australia.