Data & AI ProductOps Lead
Doha, Doha Municipality, Qatar · Full Time
Be the first to apply
- Experience
- 10+ yrs
- Salary
- —
- Openings
- 1
- Posted
- 6 days ago
- Work mode
- In office
- Education
- Bachelor's degree in Computer Science, Engineering, Data/AI or related fields
- Resume
- Required to apply
Where you'll work
Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.
Job description
Job Overview
This senior technical leadership position demands extensive practical expertise in modern data and AI infrastructure, encompassing DevOps/DataOps, production AI, observability, and integration with Service Management systems.
Key Performance Indicators
- Successful establishment, adoption, and scaling of the Data & AI ProductOps practice.
- Development of clear, reusable standards for DataOps, MLOps, LLMOps, and AgentOps.
- Implementation of ProductOps capabilities within client environments.
- Defined ownership, operational preparedness, and hierarchical L0-L3 support models.
- Enhanced product condition and compliance with Service Level Objectives (SLOs).
- Reduction in incident detection and recovery durations.
- Minimized recurring production problems and decreased unnecessary engineering escalations.
- Increased automation, proactive monitoring tools, and promotion of self-service.
- Robust integration between Data & AI ProductOps and Enterprise Service Management frameworks.
- Wider utilization of standardized patterns, playbooks, and delivery accelerators.
- Development of internal expertise and efficient client knowledge transfer.
Responsibilities
Practice Development:
- Lead and shape the Data & AI ProductOps practice, including operating models, standards, governance, reusable assets, playbooks, and implementation strategies.
- Define practice parameters across DataOps, MLOps, LLMOps, and AgentOps to operate Data, BI, ML, GenAI, and Agent-based products in production environments.
- Create reusable ProductOps resources such as operational readiness benchmarks, support frameworks, SLO guidelines, monitoring schemas, runbooks, templates, and accelerators.
Product Operations:
- Set and execute operational readiness criteria covering ownership, product criticality, support tiers, SLAs/SLOs, monitoring, alerts, escalation procedures, recovery methods, dependencies, and rollback strategies.
- Develop support models across L0 (automation and self-service), L1 (Service Desk and initial triage), L2 (ProductOps-led operational support), and L3 (complex engineering support).
- Monitor and manage product health and observability metrics such as availability, performance, data quality, pipeline integrity, model and AI performance, usage, and cost.
- Lead incident management including investigation, restoration, and root cause analysis, coordinating multiple teams and external vendors.
- Incorporate ProductOps within enterprise Incident, Problem, Change, Release, Knowledge, Service Level, and Major Incident Management processes.
- Convert repetitive incidents and operational challenges into permanent fixes, automation projects, and product enhancements.
- Institutionalize change and release management supported by automated testing, continuous integration/delivery, version control, staged deployments, and rollback plans.
- Conduct regular product operations reviews focused on health, SLO adherence, incident analysis, problem trends, technical debt, improvements, and releases.
Client Delivery:
- Evaluate client maturity in Data & AI ProductOps and develop targeted operating models, support structures, observability tools, automation, and implementation pathways.
- Design and deploy ProductOps solutions for Data, BI, ML, GenAI, and Agent products in client infrastructures.
- Lead client workshops on solution design, architecture, and execution, defining roles, responsibilities, support and operational processes, tooling, and Service Management integration.
- Oversee architecture and deployment studies for complex Data and AI environments, engaging senior stakeholders in Product, Engineering, Platform, Governance, Security, and Service Management.
- Assist in proposals, technical consulting, and advisory roles covering various ProductOps domains.
Leadership Expectations:
- Build and expand a new Data & AI ProductOps practice.
- Provide expert technical leadership while remaining actively engaged in hands-on tasks where necessary.
- Guide cross-functional teams across Data, AI, Engineering, Platform, Product, and Service Management domains.
- Deliver architecture and implementation governance on complex enterprise projects.
- Confidently engage senior clients and internal stakeholders.
- Translate emerging Data and AI technologies into practical, enterprise-ready operational standards.
- Develop internal talent and foster sustainable technical skills.
Qualifications and Experience
- Minimum bachelor’s degree in Computer Science, Engineering, Data/AI, Information Systems, or related disciplines.
- Certifications in cloud technologies, Data & AI platforms, DevOps/MLOps, Service Management, architecture, or data management are beneficial.
- At least 10 years of professional experience in Data Engineering, AI/ML, DataOps, DevOps, Site Reliability Engineering (SRE), Product Operations, Platform Engineering, or related fields.
- Extensive hands-on involvement with modern enterprise data platforms and Data/AI production solutions.
- Proven record in creating or leading production operating capabilities for Data and AI products.
- Advanced know-how in DataOps, MLOps, LLMOps, and AgentOps practices.
- Expertise in operational readiness, observability, SLA/SLO adherence, monitoring, support structures, incident and problem management, continuous integration/deployment, automated testing, release management, recovery, and rollback techniques.
- Experience managing BI and data products, pipelines, APIs, integrations, machine learning models, GenAI applications, and AI Agents in production settings.
- Practical knowledge of model lifecycle management, model deployment, drift detection, retrieval-augmented generation (RAG), vector search, AI evaluation, AI observability, agent tracing, tool execution, and operational guardrails.
- Strong familiarity with enterprise IT Service Management and integrating Data & AI teams within established support, incident, problem, change, and release workflows.
- Demonstrated skills in incident leadership, troubleshooting, root-cause analysis, and resolution of production issues.
- Experience developing technical standards, operational models, reusable frameworks, and deployment methodologies.
- Competence in client engagement, technical consulting, architectural design, solutioning, and stakeholder communication.
- Desirable experience working with government or large-scale enterprises in Qatar or the GCC region.
Preferred Technologies
- Data & Analytics Platforms: Databricks, Informatica IDMC, Microsoft Fabric, Power BI, Azure Data Services, Spark, Delta Lake.
- Data Integration & Streaming: Informatica, Kafka, APIs, batch processing, streaming, change data capture (CDC).
- AI, Machine Learning & GenAI Tools: MLflow, Azure AI/ML, model serving/monitoring, RAG, vector search, evaluation, observability.
- GenAI & Agent Frameworks: LangChain, LangGraph, Semantic Kernel or similar orchestration frameworks.
- DevOps & Platform Engineering Tools: Azure DevOps, GitHub, CI/CD pipelines, Docker, Kubernetes, Terraform, Python, SQL.
- Observability & Monitoring Platforms: OpenTelemetry, Azure Monitor, Application Insights, or equivalent tools.
- Service Management Systems: ServiceNow or comparable enterprise ITSM platforms.
Minimum education
Bachelor's Degree
Skills
How they work
Leadership
Initiative
Customer Focus
Relationship Building