- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 3 weeks ago
- Work mode
- Hybrid
- Education
- Bachelor's degree or higher in Computer Science, Mathematics, Engineering, or related field
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Nucleus AI
Nucleus AI specializes in creating and managing advanced enterprise large language models (LLMs) tailored to proprietary organizational data. Their platform delivers secure, high-performance AI solutions customized for specific workflows and maintains model control within the enterprise. They provide comprehensive AI lifecycle management including training, deployment, monitoring, and automation, enabling organizations to transition from experimentation to scalable production without infrastructure overhead. Headquartered in Singapore, the company collaborates with both enterprise and public-sector clients to seamlessly embed AI into existing data systems.
Role Description
The Tokenisation Engineer will be responsible for crafting and refining tokenization workflows suited for multilingual and specialized domain text in Nucleus AI’s LLMs. This role involves data characteristic analysis, development of vocabulary and tokenization strategies, and balancing trade-offs between model performance, computational efficiency, and accuracy. The engineer will also integrate these tokenizers into training and inference pipelines, working in close cooperation with ML engineers and researchers to experiment with and evaluate novel tokenization techniques.
Additional duties include maintaining toolsets, producing documentation, and establishing best practices to ensure tokenization processes are reliable, robust, and secure within enterprise environments. This position operates under a hybrid work model in Singapore with some remote flexibility.
Qualifications
- Solid grounding in computer science or related fields, with proven experience in NLP, language modeling, or text preprocessing.
- Practical expertise in designing and optimizing tokenization pipelines, including subword and byte-level tokenizer implementations, vocabulary construction, and managing multilingual/domain-specific datasets.
- Programming proficiency particularly in Python and relevant NLP or ML libraries, coupled with scripting and automation skills for handling data pipelines.
- Prior work with large-scale datasets, applying performance optimization techniques, and familiarity with production or research workflows involving model training and inference.
- Effective collaboration skills with cross-functional teams, ability to communicate technical concepts clearly, and maintain thorough documentation to ensure reproducibility.
- Comfortable operating within a hybrid work structure in Singapore and engaging with enterprise or government clients focusing on secure, data-centric AI systems.
- Academic credentials of a Bachelor's degree or higher in Computer Science, Mathematics, Engineering, or similar fields; experience with LLMs or enterprise AI environments is preferred.
Minimum education
Bachelor's Degree