Insight Global

Senior MLOps Engineer

Insight Global

Bengaluru, Karnataka, India · Full Time

Be the first to apply

Experience
10+ yrs
Salary
INR 1,800,000 – INR 3,300,000 / year
Openings
1
Posted
1 day ago
Work mode
In office
Education
Bachelor’s degree in Computer Science, Information Technology, or related technical discipline.
Eligibility
Any graduate is eligible to apply for this position.
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role

We are seeking a highly experienced Senior MLOps Engineer to lead the creation, automation, and governance of secure, scalable CI/CD and machine learning pipelines. This role focuses on enabling dependable, efficient, and compliant AI and software delivery across multiple platforms. You will collaborate closely with diverse teams including application development, data science, infrastructure, and security to standardize deployment processes and embed security and observability throughout the software and ML lifecycle. This position integrates DevSecOps, platform engineering, and MLOps principles with an emphasis on automation, security, and operational excellence.

Primary Responsibilities

  • Design, develop, and sustain secure and scalable CI/CD pipelines tailored for both application and AI workloads.
  • Establish and ensure adherence to best practices regarding source control, release processes, infrastructure as code, and security standards.
  • Conduct analysis of code repositories to identify security vulnerabilities and compliance gaps; lead remediation efforts.
  • Deploy, manage, and optimize container platforms using Docker and Kubernetes.
  • Develop and maintain robust and secure runtime environments across SaaS, IaaS, and PaaS.
  • Partner with platform leads and third parties to evaluate and implement system upgrades and patches.
  • Troubleshoot and resolve complex deployment and environment-related issues collaboratively with external technical teams.
  • Automate provisioning and configuration of infrastructure using Terraform, Ansible, and Azure Resource Manager.
  • Plan and carry out production deployments minimizing downtime via established release methodologies.
  • Support machine learning lifecycle stages including training, validation, deployment, monitoring, and decommissioning.
  • Construct and maintain ML pipelines leveraging tools such as MLflow, Kubeflow, or similar.
  • Oversee model version control, lifecycle tracking, and artifact management.
  • Deploy ML models as production-level APIs using REST or gRPC technologies.
  • Implement model serving frameworks including FastAPI, TorchServe, KFServing, or equivalents.
  • Design batch and real-time inference pipelines.
  • Adopt canary and blue-green deployment methods for ML model rollouts.
  • Create standardized AI deployment workflows through reusable pipeline templates and empower development and data science teams with self-service capabilities.
  • Implement observability solutions encompassing logging, monitoring, and alerting for applications and AI models.
  • Track model performance, accuracy, and detect data and concept drift.
  • Enable AI observability by capturing model inputs and outputs to support traceability and auditing.
  • Create feedback mechanisms to continuously enhance model effectiveness.
  • Secure ML pipelines and endpoints by enforcing authentication, authorization, and secrets handling.
  • Ensure responsible management of personally identifiable information (PII) and sensitive data within AI workflows.
  • Champion Responsible AI practices including governance, auditability, and risk mitigation strategies.
  • Maintain vigilance against prompt injection and security risks, particularly in large language model (LLM) environments.

Experience and Qualifications

  • Over 10 years of experience in DevSecOps, platform engineering, MLOps, or similar senior software engineering roles.
  • Proficient in continuous integration and continuous deployment tools such as Git, Bitbucket, and Jenkins.
  • Advanced knowledge of Docker, Kubernetes, and cloud-native architecture design.
  • Strong scripting skills in Python, Bash, or PowerShell.
  • Hands-on experience with major cloud platforms including AWS, Azure, or Google Cloud Platform.
  • Experience with infrastructure as code tools like Terraform and ARM; skilled in configuration management tools such as Ansible, Puppet, or Chef.
  • Familiar with code quality and security scanning tools such as Veracode and SonarQube.
  • Deep understanding of Linux systems, networking concepts, and various storage technologies including block, object, and file storage.
  • Working knowledge of CNCF ecosystem components, autoscaling techniques, and cloud observability frameworks.
  • Competence with SQL basics and messaging systems like Kafka and RabbitMQ.
  • Prior experience deploying monitoring and logging solutions for distributed systems.
  • Bachelor’s degree in Computer Science, Information Technology, or a closely related discipline.

Preferred Additional Skills

  • Master’s degree in a relevant technical field.
  • Familiarity with project and knowledge management tools such as Jira and Confluence.
  • Deployment experience with enterprise platforms like ERP, CRM, WMS, or e-commerce systems.
  • Experience supporting production environments of Large Language Model (LLM) based AI systems.

Collaboration and Work Style

  • Advocate secure development and operational best practices throughout engineering and data science teams.
  • Serve as a trusted advisor for initiatives in DevSecOps, MLOps, and platform engineering.
  • Participate actively in agile methodologies including sprint ceremonies and peer code reviews.
  • Maintain comprehensive and high-quality documentation such as architecture blueprints, technical specifications, runbooks, and procedures.
  • Promptly triage and address technical issues across development, staging, and production environments.
  • Promote open, professional communication at all organizational levels.

Minimum education

Bachelor's Degree

Tools & software

Git required Docker required Kubernetes required Amazon Web Services AWS required Microsoft PowerShell required Jenkins CI required Google Cloud Platform required Microsoft Azure required

How they work

Communication Teamwork & Collaboration Problem Solving Attention to Detail
🤖
Online · instant AI help
Broxer