S

AI Data Engineer

SG Analytics

Mumbai, Maharashtra, India · Full Time

Be the first to apply

Experience
4+ yrs
Salary
INR 2,000,000 – INR 3,500,000 / year
Openings
1
Posted
4 days ago
Work mode
In office
Education
B.Tech / B.E. in Any Specialization, Any Graduate
Eligibility
Candidates with B.Tech/B.E. in any specialization or any graduate degree are eligible to apply.
Resume
Required to apply

Where you'll work

Sign in to tell us what does and doesn't work for you here — it sharpens every match we show you.

Job description

About the Role and Company

Join SG Analytics, a part of Straive, recognized for its expertise across BFSI, Capital Markets, Technology & Media, Manufacturing, and Healthcare sectors. SG Analytics supports premier clients including Fortune 500 companies, operating globally in the U.S., U.K., Switzerland, Poland, and India. The company has been featured by Gartner, Everest Group, ISG, Deloitte Technology Fast 50 India 2024, and the Financial Times & Statista APAC 2025 High Growth Companies.

Key Responsibilities

  • Design and scale enterprise-grade lakehouse solutions utilizing Azure Databricks, ADLS Gen2, Delta Lake, and medallion architecture layers (Bronze, Silver, Gold).
  • Develop and maintain robust batch and streaming ETL/ELT pipelines with PySpark, Spark SQL, Azure Data Factory, and Databricks Workflows for production environments.
  • Construct configuration-driven, metadata-centric frameworks for efficient onboarding of diverse data formats including structured, semi-structured, and unstructured data.
  • Design and optimize gold-layer data products, vector search indexes, and embeddings leveraging Azure AI Search to enhance Retrieval-Augmented Generation (RAG) and AI agent functionalities.
  • Enforce comprehensive data governance measures including access controls and data lineage through Databricks Unity Catalog and Microsoft Purview, ensuring compliance with standards such as Section 7216, PCAOB, and NIST AI RMF.
  • Implement tuning strategies like Liquid Clustering, Z-Ordering, file compaction, and appropriately scaled compute resources to reduce platform costs.
  • Set up CI/CD pipelines, automate data quality testing, define data contracts, and establish monitoring for freshness, schema drifts, and volume changes.
  • Collaborate closely with AI developers, architects, and cross-functional teams to integrate specialized solutions into scalable platform services.

Candidate Profile

  • Minimum 4 years of practical experience building scalable data pipelines and data platform infrastructures.
  • Demonstrated expertise in designing lakehouse architectures on Microsoft Azure cloud.

Technical Skills Required

  • Proficiency with Azure cloud ecosystem including Azure Databricks, ADLS Gen2, Azure Data Factory, and Synapse Analytics.
  • Strong command of Python programming, PySpark, Spark SQL, and general SQL querying.
  • Experience with data storage formats such as Delta Lake, Parquet, Avro, and JSON.
  • Knowledge of AI-related data handling: vector search, RAG pipelines, Azure AI Search, embeddings, and chunking techniques.
  • Hands-on with governance and development operations using Databricks Unity Catalog, Microsoft Purview, CI/CD tools (Azure DevOps or GitHub Actions), and infrastructure automation (Terraform or Bicep).

Additional Information

Applicants must be ready to relocate and work onsite from Goregaon, Mumbai. The role offers an annual compensation package between 2,000,000 to 3,500,000 INR.

Eligibility and Education Criteria

The position is open to candidates holding a B.Tech, B.E. degree in any specialization or any other graduate qualification.

Minimum education

Bachelor's Degree

How they work

Teamwork & Collaboration Problem Solving Attention to Detail

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help
Broxer