Senior AI ML Architect
Hyderabad, Telangana, India (Hybrid) · Full Time
Be the first to apply
- Experience
- 10–15 yrs
- Salary
- —
- Openings
- 1
- Posted
- 1 day ago
- Work mode
- Hybrid
- Education
- Any graduate
- Eligibility
- Candidates possessing any graduate degree are eligible to apply.
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
About the Role
We are looking for a Senior AI/ML Architect based in Hyderabad with more than 10 years of technology experience. This role is a hybrid position requiring three days a week onsite presence. The successful candidate will lead the design and implementation of advanced AI and machine learning solutions, focusing on training and fine-tuning large language models (LLMs) and other deep learning models using enterprise-specific datasets. You will also develop scalable infrastructure for model hosting, serving, and lifecycle management on cloud platforms.
Responsibilities
- Design and implement AI/ML solutions for training and fine-tuning large models with domain-specific data.
- Assess and select open-weight models considering domain fit, capabilities, infrastructure needs, costs, and licensing.
- Create end-to-end pipelines covering data preparation, model training, evaluation, packaging, deployment, and ongoing monitoring.
- Hands-on fine-tuning using LoRA and QLoRA techniques, optimizing hyperparameters such as rank, learning rate, batch size, and quantization.
- Prepare and refine domain datasets through cleaning, formatting, synthetic data generation, deduplication, and quality checks.
- Utilize PyTorch, Hugging Face Transformers, PEFT, TRL, Accelerate frameworks for model training and tuning.
- Design and manage GPU-accelerated cloud training environments on AWS, Azure, or Google Cloud Platform, optimizing GPUs such as A100, H100, or L40S.
- Deploy and host models with production-grade inference frameworks like vLLM, SGLang, TensorRT-LLM, or TGI, ensuring high availability and cost-efficiency.
- Oversee the full model lifecycle including versioning, registry management, deployment, rollback, monitoring, and retirement.
- Implement MLOps pipelines through MLflow, Weights & Biases, Kubernetes, Docker, Ray, and CI/CD tools to automate training, deployment, and retraining.
- Monitor model and infrastructure metrics including performance, latency, throughput, GPU and memory utilization, and operational costs.
- Develop regression and evaluation frameworks to validate improvements over baseline models.
- Resolve issues related to model training, GPU resource management, deployment, and production environment.
- Collaborate with cross-functional teams including data scientists, ML engineers, data engineers, and cloud engineers to ensure smooth operationalization of AI/ML models.
- Establish reusable architecture patterns and engineering standards for model training and hosting tailored to domain requirements.
- Provide technical guidance and mentorship to engineering teams engaged in ML/LLM model training and deployment.
Required Expertise
- Extensive hands-on experience with PyTorch and Hugging Face Transformers for model training and fine-tuning.
- Proven background in training large language models and deep learning models using enterprise domain-specific datasets.
- Advanced knowledge of LoRA and QLoRA fine-tuning techniques.
- Experience with frameworks such as PEFT, SFT, TRL, and Accelerate.
- Familiarity with open-weight models including Llama, Qwen, Mistral, Mixtral, Gemma, DeepSeek, and Phi.
- Strong understanding of GPU utilization, training hyperparameters, dataset preparation, and model evaluation.
- Expertise in deploying and hosting ML and LLM models on AWS, Azure, or Google Cloud, including GPU infrastructure planning and management.
- Proficiency with containerization and orchestration using Docker and Kubernetes.
- Experience operating model serving frameworks like vLLM, SGLang, TensorRT-LLM, and TGI for scalable production deployments.
- Comprehensive knowledge of MLOps processes including MLflow, Weights & Biases, CI/CD automation, model versioning, and monitoring.
Qualifications
A minimum of 10 years of extensive experience in technology, particularly in AI/ML engineering, deep learning, and production-level model development.
Work Mode and Location
This position is based in Hyderabad and follows a hybrid work model requiring three days per week at the office.
Eligibility
Open to candidates with any graduate degree.
Minimum education
Bachelor's Degree