malomatia

Data & AI ProductOps Lead

malomatia

Doha, Doha Municipality, Qatar · Full Time

Be the first to apply

Experience
10+ yrs
Salary
—
Openings
1
Posted
6 days ago
Work mode
In office
Education
Bachelor's degree in Computer Science, Engineering, Data/AI or related fields
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

Job Overview

This senior technical leadership position demands extensive practical expertise in modern data and AI infrastructure, encompassing DevOps/DataOps, production AI, observability, and integration with Service Management systems.

Key Performance Indicators

  • Successful establishment, adoption, and scaling of the Data & AI ProductOps practice.
  • Development of clear, reusable standards for DataOps, MLOps, LLMOps, and AgentOps.
  • Implementation of ProductOps capabilities within client environments.
  • Defined ownership, operational preparedness, and hierarchical L0-L3 support models.
  • Enhanced product condition and compliance with Service Level Objectives (SLOs).
  • Reduction in incident detection and recovery durations.
  • Minimized recurring production problems and decreased unnecessary engineering escalations.
  • Increased automation, proactive monitoring tools, and promotion of self-service.
  • Robust integration between Data & AI ProductOps and Enterprise Service Management frameworks.
  • Wider utilization of standardized patterns, playbooks, and delivery accelerators.
  • Development of internal expertise and efficient client knowledge transfer.

Responsibilities

Practice Development:

  • Lead and shape the Data & AI ProductOps practice, including operating models, standards, governance, reusable assets, playbooks, and implementation strategies.
  • Define practice parameters across DataOps, MLOps, LLMOps, and AgentOps to operate Data, BI, ML, GenAI, and Agent-based products in production environments.
  • Create reusable ProductOps resources such as operational readiness benchmarks, support frameworks, SLO guidelines, monitoring schemas, runbooks, templates, and accelerators.

Product Operations:

  • Set and execute operational readiness criteria covering ownership, product criticality, support tiers, SLAs/SLOs, monitoring, alerts, escalation procedures, recovery methods, dependencies, and rollback strategies.
  • Develop support models across L0 (automation and self-service), L1 (Service Desk and initial triage), L2 (ProductOps-led operational support), and L3 (complex engineering support).
  • Monitor and manage product health and observability metrics such as availability, performance, data quality, pipeline integrity, model and AI performance, usage, and cost.
  • Lead incident management including investigation, restoration, and root cause analysis, coordinating multiple teams and external vendors.
  • Incorporate ProductOps within enterprise Incident, Problem, Change, Release, Knowledge, Service Level, and Major Incident Management processes.
  • Convert repetitive incidents and operational challenges into permanent fixes, automation projects, and product enhancements.
  • Institutionalize change and release management supported by automated testing, continuous integration/delivery, version control, staged deployments, and rollback plans.
  • Conduct regular product operations reviews focused on health, SLO adherence, incident analysis, problem trends, technical debt, improvements, and releases.

Client Delivery:

  • Evaluate client maturity in Data & AI ProductOps and develop targeted operating models, support structures, observability tools, automation, and implementation pathways.
  • Design and deploy ProductOps solutions for Data, BI, ML, GenAI, and Agent products in client infrastructures.
  • Lead client workshops on solution design, architecture, and execution, defining roles, responsibilities, support and operational processes, tooling, and Service Management integration.
  • Oversee architecture and deployment studies for complex Data and AI environments, engaging senior stakeholders in Product, Engineering, Platform, Governance, Security, and Service Management.
  • Assist in proposals, technical consulting, and advisory roles covering various ProductOps domains.

Leadership Expectations:

  • Build and expand a new Data & AI ProductOps practice.
  • Provide expert technical leadership while remaining actively engaged in hands-on tasks where necessary.
  • Guide cross-functional teams across Data, AI, Engineering, Platform, Product, and Service Management domains.
  • Deliver architecture and implementation governance on complex enterprise projects.
  • Confidently engage senior clients and internal stakeholders.
  • Translate emerging Data and AI technologies into practical, enterprise-ready operational standards.
  • Develop internal talent and foster sustainable technical skills.

Qualifications and Experience

  • Minimum bachelor’s degree in Computer Science, Engineering, Data/AI, Information Systems, or related disciplines.
  • Certifications in cloud technologies, Data & AI platforms, DevOps/MLOps, Service Management, architecture, or data management are beneficial.
  • At least 10 years of professional experience in Data Engineering, AI/ML, DataOps, DevOps, Site Reliability Engineering (SRE), Product Operations, Platform Engineering, or related fields.
  • Extensive hands-on involvement with modern enterprise data platforms and Data/AI production solutions.
  • Proven record in creating or leading production operating capabilities for Data and AI products.
  • Advanced know-how in DataOps, MLOps, LLMOps, and AgentOps practices.
  • Expertise in operational readiness, observability, SLA/SLO adherence, monitoring, support structures, incident and problem management, continuous integration/deployment, automated testing, release management, recovery, and rollback techniques.
  • Experience managing BI and data products, pipelines, APIs, integrations, machine learning models, GenAI applications, and AI Agents in production settings.
  • Practical knowledge of model lifecycle management, model deployment, drift detection, retrieval-augmented generation (RAG), vector search, AI evaluation, AI observability, agent tracing, tool execution, and operational guardrails.
  • Strong familiarity with enterprise IT Service Management and integrating Data & AI teams within established support, incident, problem, change, and release workflows.
  • Demonstrated skills in incident leadership, troubleshooting, root-cause analysis, and resolution of production issues.
  • Experience developing technical standards, operational models, reusable frameworks, and deployment methodologies.
  • Competence in client engagement, technical consulting, architectural design, solutioning, and stakeholder communication.
  • Desirable experience working with government or large-scale enterprises in Qatar or the GCC region.

Preferred Technologies

  • Data & Analytics Platforms: Databricks, Informatica IDMC, Microsoft Fabric, Power BI, Azure Data Services, Spark, Delta Lake.
  • Data Integration & Streaming: Informatica, Kafka, APIs, batch processing, streaming, change data capture (CDC).
  • AI, Machine Learning & GenAI Tools: MLflow, Azure AI/ML, model serving/monitoring, RAG, vector search, evaluation, observability.
  • GenAI & Agent Frameworks: LangChain, LangGraph, Semantic Kernel or similar orchestration frameworks.
  • DevOps & Platform Engineering Tools: Azure DevOps, GitHub, CI/CD pipelines, Docker, Kubernetes, Terraform, Python, SQL.
  • Observability & Monitoring Platforms: OpenTelemetry, Azure Monitor, Application Insights, or equivalent tools.
  • Service Management Systems: ServiceNow or comparable enterprise ITSM platforms.

Minimum education

Bachelor's Degree

How they work

Leadership Initiative Customer Focus Relationship Building
🤖
Online · instant AI help
Broxer