Criticalriver

Senior AI ML Architect

Criticalriver

Hyderabad, Telangana, India (Hybrid) · Full Time

Be the first to apply

Experience
10–15 yrs
Salary
—
Openings
1
Posted
1 day ago
Work mode
Hybrid
Education
Any graduate
Eligibility
Candidates possessing any graduate degree are eligible to apply.
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

We are looking for a Senior AI/ML Architect based in Hyderabad with more than 10 years of technology experience. This role is a hybrid position requiring three days a week onsite presence. The successful candidate will lead the design and implementation of advanced AI and machine learning solutions, focusing on training and fine-tuning large language models (LLMs) and other deep learning models using enterprise-specific datasets. You will also develop scalable infrastructure for model hosting, serving, and lifecycle management on cloud platforms.

Responsibilities

  • Design and implement AI/ML solutions for training and fine-tuning large models with domain-specific data.
  • Assess and select open-weight models considering domain fit, capabilities, infrastructure needs, costs, and licensing.
  • Create end-to-end pipelines covering data preparation, model training, evaluation, packaging, deployment, and ongoing monitoring.
  • Hands-on fine-tuning using LoRA and QLoRA techniques, optimizing hyperparameters such as rank, learning rate, batch size, and quantization.
  • Prepare and refine domain datasets through cleaning, formatting, synthetic data generation, deduplication, and quality checks.
  • Utilize PyTorch, Hugging Face Transformers, PEFT, TRL, Accelerate frameworks for model training and tuning.
  • Design and manage GPU-accelerated cloud training environments on AWS, Azure, or Google Cloud Platform, optimizing GPUs such as A100, H100, or L40S.
  • Deploy and host models with production-grade inference frameworks like vLLM, SGLang, TensorRT-LLM, or TGI, ensuring high availability and cost-efficiency.
  • Oversee the full model lifecycle including versioning, registry management, deployment, rollback, monitoring, and retirement.
  • Implement MLOps pipelines through MLflow, Weights & Biases, Kubernetes, Docker, Ray, and CI/CD tools to automate training, deployment, and retraining.
  • Monitor model and infrastructure metrics including performance, latency, throughput, GPU and memory utilization, and operational costs.
  • Develop regression and evaluation frameworks to validate improvements over baseline models.
  • Resolve issues related to model training, GPU resource management, deployment, and production environment.
  • Collaborate with cross-functional teams including data scientists, ML engineers, data engineers, and cloud engineers to ensure smooth operationalization of AI/ML models.
  • Establish reusable architecture patterns and engineering standards for model training and hosting tailored to domain requirements.
  • Provide technical guidance and mentorship to engineering teams engaged in ML/LLM model training and deployment.

Required Expertise

  • Extensive hands-on experience with PyTorch and Hugging Face Transformers for model training and fine-tuning.
  • Proven background in training large language models and deep learning models using enterprise domain-specific datasets.
  • Advanced knowledge of LoRA and QLoRA fine-tuning techniques.
  • Experience with frameworks such as PEFT, SFT, TRL, and Accelerate.
  • Familiarity with open-weight models including Llama, Qwen, Mistral, Mixtral, Gemma, DeepSeek, and Phi.
  • Strong understanding of GPU utilization, training hyperparameters, dataset preparation, and model evaluation.
  • Expertise in deploying and hosting ML and LLM models on AWS, Azure, or Google Cloud, including GPU infrastructure planning and management.
  • Proficiency with containerization and orchestration using Docker and Kubernetes.
  • Experience operating model serving frameworks like vLLM, SGLang, TensorRT-LLM, and TGI for scalable production deployments.
  • Comprehensive knowledge of MLOps processes including MLflow, Weights & Biases, CI/CD automation, model versioning, and monitoring.

Qualifications

A minimum of 10 years of extensive experience in technology, particularly in AI/ML engineering, deep learning, and production-level model development.

Work Mode and Location

This position is based in Hyderabad and follows a hybrid work model requiring three days per week at the office.

Eligibility

Open to candidates with any graduate degree.

Minimum education

Bachelor's Degree

Tools & software

PyTorch Docker · 10+ years required Kubernetes · 10+ years required Amazon Web Services AWS required Google Cloud Platform · 10+ years required Microsoft Azure required Hugging Face Transformers · 10+ years required

How they work

Teamwork & Collaboration Problem Solving Attention to Detail Leadership
🤖
Online · instant AI help
Broxer