- Experience
- 3+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 2 weeks ago
- Work mode
- Work from home
- Resume
- Required to apply
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About H2O.ai
Founded in 2012, H2O.ai is a leading AI company dedicated to democratizing artificial intelligence by merging Generative and Predictive AI to create custom GenAI solutions for enterprises and public agencies. The company emphasizes Sovereign AI, ensuring secure, compliant, and flexible deployments that maintain strict data privacy and control standards. Trusted by over 20,000 organizations globally, including many Fortune 500 companies, H2O.ai collaborates with notable partners such as NVIDIA, Dell Technologies, Deloitte, EY, Snowflake, AWS, and Google Cloud. The company fosters a global community of 2 million data scientists and actively supports societal causes through its AI for Good program.
Role Overview
The Senior AI Engineer role involves developing impactful AI solutions addressing complex enterprise challenges across the APAC region. This role includes designing, building, and deploying agentic AI systems, LLM-based applications, and conventional machine learning models onto production environments. Working closely with expert teams, including Kaggle Grandmasters, this position requires hands-on engineering with a strong customer interaction component and is based in Singapore.
Key Responsibilities
- Create advanced agentic AI frameworks that automate intricate, multi-step enterprise workflows.
- Develop large language model (LLM)-powered applications including retrieval-augmented generation (RAG), fine-tuning, prompt engineering, function calls, and tool integration.
- Design and deploy classical machine learning and deep learning models such as forecasting, classification, and tabular data models that effectively complement LLM solutions.
- Implement reliable guardrails, incorporate human-in-the-loop mechanisms, and establish evaluation systems to uphold production-grade AI standards.
- Manage the end-to-end AI application lifecycle from initial problem identification, data preparation, integration, testing, to deployment.
- Develop scalable backend APIs to expose AI functionalities within enterprise systems.
- Install and maintain models and services across varied environments, including cloud platforms, on-premises, and isolated air-gapped systems, addressing infrastructure needs like Kubernetes, Helm, ingress controls, TLS, and GPU support.
- Perform load testing to optimize GPU resource allocation based on real-world throughput, latency, and concurrency metrics.
- Build LLM operations infrastructure for continuous monitoring and improvement in live settings.
- Act as the primary customer liaison, liaising with data scientists, engineers, and business leaders to translate operational challenges into AI-driven solutions.
- Handle technical support queries from diagnosis through resolution.
- Lead training sessions and workshops to empower client teams in maximizing platform usage.
- Work collaboratively with internal teams including Product and Engineering to ensure polished, integrated solutions.
- Support sales initiatives and demonstrate proof-of-concept projects to build technical credibility and appropriately scope engagements.
- Efficiently manage multiple customer projects with strong organizational and prioritization skills, taking full ownership of deliverables deployed to clients.
Candidate Profile
- Minimum 3 years of practical engineering experience, with at least 1 year dedicated to agentic AI system development and deploying AI applications into production environments.
- Proven track record designing LLM-based applications involving RAG pipelines, agentic workflows, or fine-tuned models.
- Experience deploying AI models and services on cloud platforms (AWS, Azure, GCP) or enterprise settings including Kubernetes-managed on-premises environments.
- Familiarity with operating in stringent security contexts including air-gapped environments, handling offline deployments, image mirroring, private registries, and dependency packaging.
- Strong understanding of contemporary GenAI and agentic AI principles, including prompt engineering, model evaluation, guardrails, LLM operations, and multi-agent coordination protocols.
- Competence in traditional machine learning, deep learning methodologies, feature engineering, and evaluation metrics.
- Expertise in Python programming along with experience in machine learning frameworks such as PyTorch, TensorFlow, and scikit-learn; familiarity with LLM toolkits like LangChain or LlamaIndex.
- Capability to deploy open-weight models using tools like vLLM across single and multi-node environments, tuning configurations for tensor parallelism, model length, GPU memory management, and concurrency.
- Knowledge of performance trade-offs in model serving including latency, throughput, and concurrency factors.
- Working familiarity with AWS core services and GPU cost optimization.
- Skilled in Kubernetes and Helm for deployment, upgrades, and troubleshooting.
- Proficient backend developer with familiarity in REST API design, containerization (Docker/Kubernetes), and continuous integration/deployment pipelines for AI systems.
- Experience with logging and monitoring tools tailored for LLM systems, such as Fluentd, Prometheus, Grafana, and tracing frameworks like Langfuse.
- Strong problem-solving skills with the ability to navigate ambiguity while maintaining engineering rigor.
- Ability to operate under controlled change processes restricting unscheduled fixes.
- Excellent communication skills to clarify complex AI technicalities to non-technical audiences.
- Competent in preparing comprehensive technical documentation for internal and customer use such as runbooks, sizing guides, and incident reports.
Desirable Qualifications
- Experience deploying AI within regulated sectors such as public administration, finance, or healthcare.
- Previous customer-facing engineering roles with active client engagement.
- Up-to-date knowledge of advancements in AI and agentic AI with ability to integrate innovations into projects.
- Understanding of enterprise security enhancements like mutual TLS, log redaction, secrets management, role-based access control, and auditing.
- Participation in Kaggle competitions or similar competitive machine learning platforms.
Why Join H2O.ai?
- Leading market rewards
- Culture supportive of remote work
- Flexible working arrangements
- Opportunity to work with an elite, world-class team
- Focus on career advancement
Diversity & Inclusion
H2O.ai is dedicated to fostering an inclusive work environment. Employment opportunities are offered without discrimination based on race, ethnicity, religion, gender, sexual orientation, age, disability, or any legally protected characteristic.
Industry
Artificial Intelligence