Building enterprise-scale data pipelines at Walmart — transforming raw data into reliable, queryable assets that power business decisions.
Data Engineer with 3+ years of experience building large-scale data pipelines on GCP. Currently owning the consumption layer of Walmart's enterprise data platform — Spark-Scala, Airflow, BigQuery across 100+ SAP and Kafka source tables.
M.Tech from ISI Kolkata (AIR 11) with research on Graph Transformers. Previously built backend APIs, ML models, and automated reporting at Tata Power. Also ships full-stack apps — deployed a production Next.js app with AI-powered writing.
My professional journey in data engineering and software development.
Strong foundation in computer science and data science.
Technologies and tools I work with daily.
Highlights from my professional experience.
Owning consumption layer in Spark-Scala, ingestion from 100+ SAP tables and Kafka brokers into GCS/BigQuery. Airflow DAGs with SLA monitoring and automated failure recovery.
Novel Graph Transformer replacing O(N²) global attention with partition-aware attention. Outperformed SGFormer on 10/12 benchmarks. Scales to 1.6M+ node graphs.
AI-powered note-taking app with rich editor, code blocks, math rendering, drawing canvas, PDF export. Integrates Gemini and Groq for AI writing. Deployed as PWA.
Python tool that standardized DAG and CCM configuration across new pipeline onboardings — column sensitivity detection, schema mapping, GCS provisioning.
ML forecasting model for day-ahead electricity trading on Indian Energy Exchange, achieving 88% accuracy with real-time PI Server integration.
End-to-end automated Power BI dashboards for HR, customer operations, and engineering teams. REST APIs with FastAPI and Node.js serving 10+ dashboards.
Optimized XGBoost model improving payment default prediction from 77% to 85% accuracy using geographic feature engineering and hyperparameter tuning.