Research Engineer, Post-Training
Bengaluru, Karnataka, India · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 1 week ago
- Work mode
- In office
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About Jurisphere
Jurisphere is pioneering AI systems designed to execute complex legal tasks. Having secured a $2.2M seed funding led by Info Edge Ventures with additional support from Flourish Ventures, Antler, and 8i Ventures, Jurisphere's platform is currently utilized by over 500 organizations and more than 20,000 legal professionals, including notable names such as ICICI Bank, Unilever, Philips, Tata Capital, and CMS IndusLaw. Our compact team innovates at the intersection of advanced AI, legal reasoning, and critical professional processes.
Role Overview
We are seeking a Research Engineer with a focus on post-training to enhance our models’ proficiency in legal tasks. This role involves hands-on research engineering, working extensively with open-weight models and addressing complex, ambiguous challenges.
Key Responsibilities
- Conduct post-training experiments across techniques such as Supervised Fine-Tuning (SFT), preference optimization, Reinforcement Learning with Human Feedback (RLHF/RLAIF), reward modelling, and model distillation.
- Develop training datasets derived from expert feedback, model activity logs, legal workflows, and generated synthetic data.
- Create evaluations and reward frameworks to measure legal reasoning, research accuracy, drafting quality, document analysis, and citation precision.
- Analyze model and agent behaviors to detect and understand consistent failure patterns.
- Develop agent environments that incorporate retrieval systems, auxiliary tools, subagents, validation processes, and workflows requiring extended decision-making horizons.
- Translate research insights into improved training data, model architectures, evaluation metrics, and production-level systems.
- Collaborate closely with engineering teams and legal experts to convert expert assessments into scalable, reliable systems.
Required Qualifications and Skills
- Practical experience in training or post-training of Large Language Models (LLMs), especially open-weight models.
- Proficient in Python programming and machine learning engineering.
- Experience with a range of methods including Supervised Fine-Tuning, preference learning, reinforcement learning, reward modelling, distillation, agent frameworks, or evaluation methodologies.
- Strong capability in experimental design—forming hypotheses, conducting controlled tests, analyzing model outputs, and interpreting inconclusive outcomes.
- Capability to independently manage technically challenging and loosely defined problems.
Preferred but Not Required
- Hands-on knowledge of distributed training or GPU acceleration.
- Experience developing agent-based systems and evaluation infrastructure.
- Contributions to research publications or open-source projects.
- Familiarity with critical high-stakes domains such as legal, financial, or healthcare sectors.
Role Significance
Within Jurisphere, post-training represents a critical feedback cycle involving models, legal practice, expert input, and technical refinement. This position offers the opportunity to oversee the entire improvement loop: identifying model shortcomings, designing effective evaluations, crafting training improvements, and verifying enhanced performance. The ultimate objective is to develop AI models that considerably improve the execution of legal work.