About
Data Engineer with 6 years of experience building scalable batch and real-time data pipelines and data architectures across on-premise and cloud environments. Specialized in Spark, PySpark, Databricks, and lakehouse architectures, with experience delivering data platforms, optimizing pipeline performance, and enabling data-driven products.
Experience
-
Data EngineerPolestar GlobalJul 2024 – Sep 2026
Acted as Data Product Owner and Lead Data Engineer for Voyages & Routes and Maritime Transparency Index (MTI). Led development of scalable batch and near real-time pipelines, geospatial analytics, Apache Iceberg operators, source-priority based MDM, and data quality frameworks. Migrated legacy ETL workloads into AWS.
-
Data EngineerImpetus TechnologiesSep 2022 – Jun 2024
Migrated legacy Mainframe and Datastage pipelines into Databricks using Delta Live Tables. Built scalable pipelines using Medallion architecture, implemented DLT-META framework, migrated enterprise systems from Vertica to Hadoop/Spark, and developed validation and reconciliation frameworks.
-
Big Data EngineerProjectPro TechnologiesApr 2022 – Jul 2022
Worked on technologies including PySpark, Apache Spark, AWS S3, AWS EMR, AWS Athena, AWS EC2, Docker, SQL, and data warehousing.
-
Big Data EngineerAxis BankAug 2020 – Apr 2022
Developed batch and real-time pipelines on Cloudera Hadoop ecosystem. Built ETL workflows using Spark, Hive, Sqoop and Oozie. Integrated Kafka streaming for near real-time ingestion and optimized Hive and Impala query performance for analytics workloads.
Education
-
PG DiplomaCentre for Development of Advanced Computing2019 – 2020
-
B.Tech / B.E.N.B.N. Sinhgad College of Engineering, Solapur2014 – 2017
Skills
Tools / apps / platforms
Languages
Projects
-
Research Topic Recommendation SystemMongoDB, Python, pandas, flask, sklearn, HTML
Recommendation system to find relevant research topics using aggregation framework and FP Growth, deployed on cloud as a service.
-
Data Warehouse to Data Lake MigrationVertica, Hadoop, Spark
Migrated enterprise systems from Vertica to Hadoop/Spark and developed validation and reconciliation frameworks.
-
Mainframe and Datastage Migrations to DatabricksDatabricks, Delta Live Tables, Medallion architecture
Migrated legacy Mainframe and Datastage pipelines into Databricks using Delta Live Tables and built standardized pipeline development frameworks.
-
Lakehouse Platform with IcebergApache Iceberg, AWS, Airflow
Built data lakehouse architecture on AWS and developed Airflow operators for managing Iceberg tables and pipelines.
-
Maritime Risk Scoring Platform (MTI)
Designed end-to-end data pipeline for ship risk scoring using multiple data sources and enabled real-time analytics for maritime intelligence.
-
AI-Powered Data Platform ToolingLangdock, Claude Desktop, MCP, AWS Athena
Built LLM-powered metadata exploration and SQL generation agents, metadata enrichment automation, and natural language querying POC to improve self-serve data exploration.
-
Voyages & Routes Database
Create a voyage and route database that segments transits at meaningful boundaries and aggregates them into economic voyages for emissions scoring, trade flow analysis, and voyage prediction.
Courses & certifications
- Security and Privacy for Big Data
🎯 Hobbies & interests
- Hiking
- Mythology