Sripad Pujari

Sripad Pujari

Data Engineer · Spark, PySpark, Databricks & Lakehouse Architectures

Solapur, Maharashtra, India

@sripad_pujari

0 followers

About

Data Engineer with 6 years of experience building scalable batch and real-time data pipelines and data architectures across on-premise and cloud environments. Specialized in Spark, PySpark, Databricks, and lakehouse architectures, with experience delivering data platforms, optimizing pipeline performance, and enabling data-driven products.

Experience

  • Data Engineer
    Polestar Global
    Jul 2024 – Sep 2026

    Acted as Data Product Owner and Lead Data Engineer for Voyages & Routes and Maritime Transparency Index (MTI). Led development of scalable batch and near real-time pipelines, geospatial analytics, Apache Iceberg operators, source-priority based MDM, and data quality frameworks. Migrated legacy ETL workloads into AWS.

  • Data Engineer
    Impetus Technologies
    Sep 2022 – Jun 2024

    Migrated legacy Mainframe and Datastage pipelines into Databricks using Delta Live Tables. Built scalable pipelines using Medallion architecture, implemented DLT-META framework, migrated enterprise systems from Vertica to Hadoop/Spark, and developed validation and reconciliation frameworks.

  • Big Data Engineer
    ProjectPro Technologies
    Apr 2022 – Jul 2022

    Worked on technologies including PySpark, Apache Spark, AWS S3, AWS EMR, AWS Athena, AWS EC2, Docker, SQL, and data warehousing.

  • Big Data Engineer
    Axis Bank
    Aug 2020 – Apr 2022

    Developed batch and real-time pipelines on Cloudera Hadoop ecosystem. Built ETL workflows using Spark, Hive, Sqoop and Oozie. Integrated Kafka streaming for near real-time ingestion and optimized Hive and Impala query performance for analytics workloads.

Education

  • PG Diploma
    Centre for Development of Advanced Computing
    2019 – 2020
  • B.Tech / B.E.
    N.B.N. Sinhgad College of Engineering, Solapur
    2014 – 2017

Skills

2 to 5 years 2 to 5 years Several years of experience.
1 to 2 years 1 to 2 years A year or two of steady use.
3 to 6 months 3 to 6 months A few months of practice with it.

Tools / apps / platforms

2 to 5 years 2 to 5 years Several years of experience.
Claude Apache Hadoop Apache Airflow GitHub Copilot Amazon Web Services AWS
1 to 2 years 1 to 2 years A year or two of steady use.
Databricks Apache Hive Apache Kafka
3 to 6 months 3 to 6 months A few months of practice with it.
AWS S3 Docker
<3 months <3 months Just getting started - under three months. This is also what shows when a level has not been set.
MongoDB

Languages

Fluent Fluent Comfortable using it for work.
Hindi Speak English Speak Marathi Speak

Projects

  • Research Topic Recommendation System
    MongoDB, Python, pandas, flask, sklearn, HTML

    Recommendation system to find relevant research topics using aggregation framework and FP Growth, deployed on cloud as a service.

  • Data Warehouse to Data Lake Migration
    Vertica, Hadoop, Spark

    Migrated enterprise systems from Vertica to Hadoop/Spark and developed validation and reconciliation frameworks.

  • Mainframe and Datastage Migrations to Databricks
    Databricks, Delta Live Tables, Medallion architecture

    Migrated legacy Mainframe and Datastage pipelines into Databricks using Delta Live Tables and built standardized pipeline development frameworks.

  • Lakehouse Platform with Iceberg
    Apache Iceberg, AWS, Airflow

    Built data lakehouse architecture on AWS and developed Airflow operators for managing Iceberg tables and pipelines.

  • Maritime Risk Scoring Platform (MTI)

    Designed end-to-end data pipeline for ship risk scoring using multiple data sources and enabled real-time analytics for maritime intelligence.

  • AI-Powered Data Platform Tooling
    Langdock, Claude Desktop, MCP, AWS Athena

    Built LLM-powered metadata exploration and SQL generation agents, metadata enrichment automation, and natural language querying POC to improve self-serve data exploration.

  • Voyages & Routes Database

    Create a voyage and route database that segments transits at meaningful boundaries and aggregates them into economic voyages for emissions scoring, trade flow analysis, and voyage prediction.

Courses & certifications

  • Security and Privacy for Big Data

🎯 Hobbies & interests

  • Hiking
  • Mythology

Explore the Broxer community

🤖
Online · instant AI help
Broxer