Open for Select Freelance & Architecture Advisory Roles: High-scale Spark tuning, Medallion Lakehouses & domain data platforms. Inquire Availability
Biography & Career Journey

About Sagar D. Shingare

Transforming petabyte-scale data complexity into resilient, high-speed, and cost-efficient distributed platforms.

Sagar D. Shingare
Big Data Tech Lead & Architect Global IT & FinTech Practice

Quick Facts

  • 14+ Years in Tech & Data Systems
  • Based in Pune, Maharashtra, India
  • M.Sc & B.Sc in Computer Science
  • English, Hindi, Marathi
  • Spark (Scala/Python), Delta Lake, Databricks

Engineering Data at Scale

I am an enterprise Data Architect and Big Data Tech Lead with over 14 years of total software engineering experience and 8+ years specializing in distributed systems, modern cloud lakehouses, and high-throughput real-time pipelines.

Currently leading enterprise BFSI data engineering initiatives as a Big Data Tech Lead & Data Architect, I design and oversee cloud-native analytics platforms for premier global banking and financial services institutions. My work ensures high-throughput transactional reconciliation, sub-second fraud detection, and automated regulatory reporting under stringent SLAs.

Previously as an enterprise Data Architect, I spearheaded multi-terabyte data lakehouse implementations across Life Sciences, Retail, and Insurance clients—standardizing reusable data engineering frameworks and slashing batch runtimes by over 60%. As a Big Data Developer, I built high-volume promotional AI and oncology analytics pipelines using PySpark and Databricks, and across premier enterprise consulting engagements, I developed large-scale product data lake ingestion systems handling 10TB+ datasets.

Core Engineering Principles

1. Idempotency & Fault Tolerance

Pipelines must be self-healing and deterministic. Re-running a batch or stream slice must never produce duplicates or corrupt downstream state.

2. Medallion Layer Rigor

Strict separation between raw ingest (Bronze), validated & normalized CDC (Silver), and business-curated marts (Gold) prevents data swamp degradation.

3. Cloud FinOps by Design

Data volume growth shouldn't explode cloud bills. Partition pruning, broadcast joins, and right-sized ephemeral compute keep costs under control.

4. Open Formats Over Lock-In

Standardizing on Delta Lake and Apache Iceberg ensures seamless query interoperability across Spark, Trino, BigQuery, and Snowflake.

Track Record

Career Journey & Leadership Roles

A continuous trajectory of delivering scalable data systems across Tier-1 organizations.

Senior Big Data Tech Lead & Data Architect

Global IT & Financial Services Practice (BFSI)
Jan 2026 – Present Pune, India

Leading enterprise BFSI data engineering delivery, architecting cloud-native ETL and scalable analytics platforms for premier banking and financial services clients. Driving high-throughput data pipelines, real-time transaction streaming, fraud detection pipelines, and regulatory reporting platforms under stringent SLAs.

Apache Spark PySpark AWS BigQuery Kafka Airflow Python SQL Banking & Financial Platforms

Data Architect

Enterprise Data & AI Practice
Jan 2022 – Dec 2025 Pune, India (Remote)

Architected and implemented enterprise-grade data engineering solutions for Life Sciences, Retail, Insurance, Consumer Tech, and AI/ML programs. Designed 3TB+ daily scalable ingestion, ELT/ETL, and Medallion Lakehouse transformations using Spark, Databricks, BigQuery, and cloud orchestration tools. Developed standardized, reusable data engineering frameworks.

Spark Scala Python Databricks BigQuery AWS EMR Delta Lake Apache Iceberg Medallion Architecture Airflow

Big Data Developer

Global Digital & Technology Services
Jan 2020 – Dec 2021 Pune, India

Delivered high-volume data workflows for CPG forecasting and Life Sciences analytics using Azure Databricks, AWS EMR, and Cloudera. Built ingestion and analytics pipelines for a global Sales & Marketing Promotional AI platform using PySpark. Integrated 26+ diverse Life Sciences data sources for Oncology analytics and automated ETL pipelines with Python.

PySpark Azure Databricks AWS EMR Cloudera Hadoop Hive Python Teradata Migration

Big Data Technology Specialist

Enterprise Technology Consulting
Apr 2014 – Dec 2019 Bangalore / Pune, India

Engineered core components of large-scale enterprise product data lakes using Hadoop ecosystem technologies (Spark, Hive, Sqoop). Managed distributed cluster performance, validated and harmonized 10TB+ ingested data feeds, and migrated relational Oracle workloads to distributed platforms.

Hadoop Spark Hive Sqoop MapReduce Oracle SQL Server PostgreSQL Linux Shell Scripting

Software Developer

EdTech Enterprise Solutions
Jan 2011 – Feb 2014 Pune, India

Delivered backend systems, mobile applications, and automation solutions for educational institutions. Enhanced enterprise student information systems (SIS) database reliability and optimized relational data operations across MySQL and Linux server environments. Built shell scripts for automated operations and deployed customized solutions across 30+ institutions.

Core Java Android MySQL Linux Shell Scripting JavaScript
Foundations

Academic Credentials

Master of Computer Science (M.Sc)
Pune University
2010 – 2013 • First Class

Specialized in advanced database systems, algorithms, distributed computing architectures, and software engineering principles.

Bachelor of Computer Science (B.Sc)
Solapur University (Sangola College)
2006 – 2009 • Distinction

Foundation in programming languages (Java, C, C++), relational database design, data structures, and operating systems.

Let's Collaborate

Want to discuss an enterprise architecture or freelance project?

Reach out directly for consulting inquiries, technical advisory, or speaking opportunities.