- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 6 days ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Chipforge
Chipforge is transforming the engineering workflow by converting design ideas expressed in plain English directly into synthesizable, verified RTL code. Beginning with FPGAs and moving towards ASICs, their solution eliminates the delays, repetitive iterations, and complex tooling traditionally slowing this process. Their systems operate in various settings, from shared cloud infrastructure to fully isolated air-gapped environments to meet strict security and data residency mandates.
Role Overview
This position is responsible for managing and curating the datasets used for model training at Chipforge. The Data Research Engineer will determine inclusion criteria for training data, develop pipelines to assemble datasets, and establish automated validation procedures to ensure data quality in a demanding domain where verifying correctness is challenging and costly.
Key Responsibilities
- Oversee the complete lifecycle of the training dataset—from sourcing data, filtering for appropriate licenses, eliminating duplicates, preparing datasets, to managing version control enabling reproducible training runs.
- Create synthetic data integrated with verification steps; generate training examples rigorously validated against actual tools instead of relying solely on model outputs, retaining only verified samples.
- Design, calibrate, and maintain automated quality control filters to uphold the standards of data inclusion and promptly detect quality drift to prevent substandard data from entering training pipelines.
- Handle data sources beyond publicly available code, including data generated by Chipforge's own tools, ensuring strict adherence to privacy and data-residency requirements of client environments.
- Implement contamination prevention mechanisms and construct held-out datasets to guarantee the integrity and validity of evaluation results.
- Convert expertise from hardware engineers regarding correctness criteria into automated code implementations scalable beyond manual review capabilities.
- Drive the data development roadmap in collaboration with the Head of Product Engineering, aligning dataset priorities with evolving product requirements.
Qualifications and Experience Needed
- Previous experience designing synthetic training datasets with built-in automated verification, not just collecting data via prompts; understanding generation, checking, filtering processes, with clear evidence supporting the data's quality.
- Keen judgment on data quality with the ability to rationalize decisions on retaining or discarding individual data examples.
- Strong Python programming skills geared towards production environments and pipeline engineering, familiar with object storage systems, dataset versioning, orchestration tools, and ensuring pipeline determinism.
- Solid understanding of the path training data takes to influence model performance, including necessary adaptations when models evolve.
- Expertise in deduplication techniques and contamination controls such as near-duplicate detection and held-out data overlap checks.
- Knowledge of open-source licensing considerations related to training datasets.
- Ability to read and comprehend code in unfamiliar programming languages, especially Verilog and SystemVerilog, sufficient to identify faulty samples and communicate effectively with hardware engineers.
- Familiarity with large language models (LLMs), their failure modes, and strategies to build dependable systems around inherently non-deterministic components.
- Comfort working in ambiguous situations, making informed decisions without established protocols, and constructively defending those decisions.
Desirable Attributes
- Experience with reinforcement learning techniques based on verifiable rewards, rejection sampling, or preference data collection.
- Hands-on involvement in continued model pretraining or fine-tuning, including knowledge of distributed training data formats.
- Exposure to hardware design processes or EDA tools, such as Verilog simulation or synthesis; prior hardware expertise is not obligatory but interest is valued.
- Experience in evaluation and benchmarking, particularly related to code generation capabilities.
- Background in developing software solutions for air-gapped or on-premises deployment settings.
- Previous mentoring or leadership experience, as this role may evolve to include mentoring responsibilities.
Working Conditions
This is a full-time remote-first position open to applicants based anywhere in Australia. The team collaborates closely across geographic locations including Australia, Singapore, and India. Candidates should be comfortable working across multiple time zones. While the role is predominantly remote, occasional travel for team or product meetings may be expected.