About Sagar D. Shingare
Transforming petabyte-scale data complexity into resilient, high-speed, and cost-efficient distributed platforms.
Quick Facts
- 14+ Years in Tech & Data Systems
- Based in Pune, Maharashtra, India
- M.Sc & B.Sc in Computer Science
- English, Hindi, Marathi
- Spark (Scala/Python), Delta Lake, Databricks
Engineering Data at Scale
I am an enterprise Data Architect and Big Data Tech Lead with over 14 years of total software engineering experience and 8+ years specializing in distributed systems, modern cloud lakehouses, and high-throughput real-time pipelines.
Currently leading enterprise BFSI data engineering initiatives as a Big Data Tech Lead & Data Architect, I design and oversee cloud-native analytics platforms for premier global banking and financial services institutions. My work ensures high-throughput transactional reconciliation, sub-second fraud detection, and automated regulatory reporting under stringent SLAs.
Previously as an enterprise Data Architect, I spearheaded multi-terabyte data lakehouse implementations across Life Sciences, Retail, and Insurance clients—standardizing reusable data engineering frameworks and slashing batch runtimes by over 60%. As a Big Data Developer, I built high-volume promotional AI and oncology analytics pipelines using PySpark and Databricks, and across premier enterprise consulting engagements, I developed large-scale product data lake ingestion systems handling 10TB+ datasets.
Core Engineering Principles
1. Idempotency & Fault Tolerance
Pipelines must be self-healing and deterministic. Re-running a batch or stream slice must never produce duplicates or corrupt downstream state.
2. Medallion Layer Rigor
Strict separation between raw ingest (Bronze), validated & normalized CDC (Silver), and business-curated marts (Gold) prevents data swamp degradation.
3. Cloud FinOps by Design
Data volume growth shouldn't explode cloud bills. Partition pruning, broadcast joins, and right-sized ephemeral compute keep costs under control.
4. Open Formats Over Lock-In
Standardizing on Delta Lake and Apache Iceberg ensures seamless query interoperability across Spark, Trino, BigQuery, and Snowflake.
Career Journey & Leadership Roles
A continuous trajectory of delivering scalable data systems across Tier-1 organizations.
Senior Big Data Tech Lead & Data Architect
Global IT & Financial Services Practice (BFSI)Leading enterprise BFSI data engineering delivery, architecting cloud-native ETL and scalable analytics platforms for premier banking and financial services clients. Driving high-throughput data pipelines, real-time transaction streaming, fraud detection pipelines, and regulatory reporting platforms under stringent SLAs.
Data Architect
Enterprise Data & AI PracticeArchitected and implemented enterprise-grade data engineering solutions for Life Sciences, Retail, Insurance, Consumer Tech, and AI/ML programs. Designed 3TB+ daily scalable ingestion, ELT/ETL, and Medallion Lakehouse transformations using Spark, Databricks, BigQuery, and cloud orchestration tools. Developed standardized, reusable data engineering frameworks.
Big Data Developer
Global Digital & Technology ServicesDelivered high-volume data workflows for CPG forecasting and Life Sciences analytics using Azure Databricks, AWS EMR, and Cloudera. Built ingestion and analytics pipelines for a global Sales & Marketing Promotional AI platform using PySpark. Integrated 26+ diverse Life Sciences data sources for Oncology analytics and automated ETL pipelines with Python.
Big Data Technology Specialist
Enterprise Technology ConsultingEngineered core components of large-scale enterprise product data lakes using Hadoop ecosystem technologies (Spark, Hive, Sqoop). Managed distributed cluster performance, validated and harmonized 10TB+ ingested data feeds, and migrated relational Oracle workloads to distributed platforms.
Software Developer
EdTech Enterprise SolutionsDelivered backend systems, mobile applications, and automation solutions for educational institutions. Enhanced enterprise student information systems (SIS) database reliability and optimized relational data operations across MySQL and Linux server environments. Built shell scripts for automated operations and deployed customized solutions across 30+ institutions.
Academic Credentials
Specialized in advanced database systems, algorithms, distributed computing architectures, and software engineering principles.
Foundation in programming languages (Java, C, C++), relational database design, data structures, and operating systems.