- Experience
- 3+ yrs
- Salary
- USD 30 – USD 50 / hour
- Openings
- 1
- Posted
- 3 days ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
We are seeking an AI Prompt Engineer to join our team remotely in India. The role involves designing, testing, and refining prompts and evaluation processes that enhance the performance of large language models (LLMs) within real-world production systems. You will partner closely with both engineers and researchers to develop prompt libraries, conduct A/B testing, establish evaluation rubrics, and boost the dependability, accuracy, and safety of AI assistants, copilots, agents, and retrieval-augmented generation (RAG) systems.
Key Responsibilities
- Manage the entire prompt lifecycle including crafting, debugging, version control, and maintaining prompt libraries.
- Develop comprehensive prompt evaluation frameworks with rubrics assessing accuracy, completeness, adherence to instructions, tone, and compliance with policies.
- Generate reinforcement learning from human feedback (RLHF)-style preference and ranking datasets with clearly articulated rationales to facilitate instruction tuning.
- Translate product requirements into system-wide prompts, developer prompts, and standardized user prompt patterns.
- Enhance model grounding and minimize hallucinations by employing RAG methodologies and enforcing citation and traceability rules.
- Author annotation guidelines and assist quality assurance evaluations addressing edge cases, escalation procedures, and consistency checks.
- Track quality indicators such as pass rates, inter-annotator agreement scores, and regression failures throughout release cycles.
- Contribute to safety alignment initiatives through red-teaming of prompts, enhancing jailbreak resistance, and preventing sensitive data leaks.
Required Qualifications
- Minimum of three years’ experience in software engineering, machine learning or natural language processing workflows, or practical LLM prompt engineering.
- Proficiency in constructing clear, structured instructions with complex multi-step constraints.
- Hands-on experience performing LLM evaluations using rubrics, pairwise ranking systems, golden data sets, and adversarial testing techniques.
- Understanding of RLHF approaches and human-in-the-loop feedback mechanisms.
- Familiarity with Python programming and experimentation tools such as notebooks, scripts, JSON handling, and APIs.
Preferred Qualifications
- Experience working with RAG pipelines, vector-based search, evaluation of embeddings, and grounded answer verification.
- Knowledge about content safety labeling systems, policy taxonomies, and implementation of prompt-based content guardrails.
- Exposure to tools for LLM observability and evaluation like regression test suites and evaluation harnesses.
Compensation
The position offers a competitive hourly wage ranging from $30 to $50.
Additional Information
This is a fully remote position targeted specifically at candidates based in India. The role is focused on distributed collaboration across teams.