- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 6 days ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Chipforge
Chipforge transforms engineers' straightforward design intentions into verified, synthesizable RTL, initially focusing on FPGAs and planning to extend to ASICs. This approach significantly reduces the time, cost of iterations, and complexity of tools typically involved. Chipforge supports diverse deployment environments, from shared cloud platforms to fully isolated setups that meet stringent security and data residency standards.
Role Overview
This position is responsible for managing the data underpinning our model operations. Your duties include deciding which data is included in training sets, constructing the pipelines to produce these datasets, and establishing validation methods to ensure training data quality despite the high cost of correctness verification.
You will collaborate closely with the training engineer, hardware experts who validate design correctness, and the evaluation team. Although hardware expertise isn't mandatory, familiarity with digital design and Verilog/VHDL is expected to effectively translate expert knowledge into data pipelines.
Key Responsibilities
- Manage the entire training corpus process: sourcing, license verification, deduplication, data preparation, and creating reproducible, version-controlled datasets.
- Generate synthetic data incorporating verification processes, utilizing real tooling validation rather than model predictions to retain only verified examples.
- Define and maintain the quality standards for training data, creating automated filtering systems to prevent poor data and drift from affecting training runs.
- Handle diverse data sources beyond open code, including signals from internal tools, while respecting privacy and data residency constraints imposed by customer environments.
- Implement contamination control and held-out data construction to ensure genuine and reliable evaluation results.
- Convert hardware expert judgments into executable code, encoding correctness criteria that surpass the scale manageable by human reviewers.
- Develop the data roadmap collaboratively with the Head of Product Engineering to align with product launch requirements.
Required Qualifications and Experience
- Proven experience creating synthetic training data integrated with automated verification loops, including generation, validation, filtering, and providing defensible quality assurance.
- Strong discernment regarding data quality, with the ability to justify data selection decisions based on real examples.
- Advanced Python programming skills and experience in pipeline engineering including object storage management, dataset versioning, orchestration, and ensuring deterministic job outputs.
- Understanding of the entire training data lifecycle and adaptability to model evolution.
- Expertise in deduplication techniques, contamination control, near-duplicate detection, and maintaining held-out data sets.
- Knowledgeable in open-source license compliance as it pertains to training datasets.
- Ability to read and comprehend code written in languages outside your primary expertise, especially Verilog and SystemVerilog, to evaluate sample integrity and communicate effectively with hardware engineers.
- Familiarity with large language models (LLMs), their behavior, limitations, and strategies to build dependable systems around non-deterministic components.
- Comfortable making decisions with ambiguous guidelines and defending these calls.
Preferred Skills
- Experience with reinforcement learning approaches such as verifiable reward systems, rejection sampling, or preference data collection.
- Hands-on knowledge of continual pretraining or fine-tuning methods, including working with distributed training data formats.
- Interest or exposure to hardware design processes or electronic design automation (EDA) tools like Verilog simulation or synthesis.
- Background in evaluation and benchmarking, especially for code generation tasks.
- Experience developing solutions for air-gapped or on-premises deployment scenarios.
- Previous mentoring experience, as this role may expand to include leadership responsibilities.
Work Environment
This position is primarily remote, open to applicants based anywhere in Australia. The team operates across Australia, Singapore, and India, necessitating comfort with cross-time-zone collaboration. Occasional travel for team or product meetings may be required but day-to-day work is remote.