Data Engineer with a passion
for building at scale

Ranjan Kumar Choubey

Data Engineer with 3+ years of experience building large-scale data pipelines on GCP. Currently owning the consumption layer of Walmart's enterprise data platform — Spark-Scala, Airflow, BigQuery across 100+ SAP and Kafka source tables.

M.Tech from ISI Kolkata (AIR 11) with research on Graph Transformers. Previously built backend APIs, ML models, and automated reporting at Tata Power. Also ships full-stack apps — deployed a production Next.js app with AI-powered writing.

3+ Years Experience
AIR 11 ISI Entrance Exam
100+ Tables Ingested
20+ Airflow DAGs

Where I've worked

My professional journey in data engineering and software development.

Consultant — Data Engineer Deloitte
May 2025 – Present
Client: Walmart — Enterprise Data Platform (Ingestion & Consumption Layer)
  • Own the consumption layer — writing Spark-Scala transformation pipelines, building JARs, end-to-end data flow from Datalake to BigQuery.
  • Delivered the UEP Sales Module independently: requirements gathering, SQL validation against SAP, Spark-Scala development, Airflow orchestration, and SFTP delivery.
  • Built ingestion pipelines for 100+ tables across SAP-ECC, EWM, CAR; also ingested data from Kafka brokers into GCS Datalake.
  • Designed Airflow DAGs orchestrating Spark jobs and SFTP transfers with retry policies, SLA monitoring, and automated failure recovery.
  • Implemented data quality validation using SQL checks against SAP sources. Debugged and resolved production pipeline failures.
  • Built a Python automation tool that standardized DAG and CCM configuration across new pipeline onboardings.
GATE DA Faculty (Part-time) IMS Academy
2025 – Present
  • Teaching Probability & Statistics and Machine Learning for GATE Data Science & AI aspirants on weekends.
Data Science Intern Allstate India
May 2024 – Jul 2024
  • Optimized baseline XGBoost model for payment default prediction, improving accuracy from 77% to 85% via geographic feature engineering.
  • Analyzed geographic data to identify default trends; presented to Bengaluru & US teams, leading to 3 targeted model enhancements.
Lead Engineer Tata Power
Sep 2022 – Aug 2023
  • Created ML forecasting model for day-ahead electricity trading on IEX, achieving 88% prediction accuracy.
  • Engineered automated pipeline integrating real-time plant telemetry (PI Server) with IEX market data, reducing manual collection by 50%.
  • Designed backend REST APIs using FastAPI and Node.js to serve data for 10+ Power BI dashboards.
Graduate Engineer Trainee Tata Power
Aug 2021 – Sep 2022
  • Deployed ML models on on-premise servers, managing full lifecycle from development to production with CI/CD.
  • Built email automation for employee onboarding — extracting, parsing, and updating DB, eliminating manual entry for 100+ employees/year.
  • Automated Power BI reporting for HR, operations, and engineering teams, reducing manual effort by 20%.

Academic background

Strong foundation in computer science and data science.

M.Tech — Computer Science (Data Science)
Indian Statistical Institute, Kolkata 2023 – 2025
AIR 11 in ISI Entrance AIR 1 EWS Category
B.Tech — Information Technology
BIT Sindri, Dhanbad 2017 – 2021
CGPA 8.89 — Rank 1 BITSAA Scholar '18 & '19

Tech stack

Technologies and tools I work with daily.

Languages
Python Scala SQL Bash
Data Engineering
Apache Airflow Apache Spark Kafka ETL/ELT BigQuery
Cloud & Storage
GCP GCS Dataproc Firebase
Databases
PostgreSQL MySQL SQL Server
Backend & APIs
FastAPI Node.js REST APIs
DevOps & Tools
Docker Git CI/CD Linux Power BI
ML & Analytics
Machine Learning Deep Learning PyTorch NLP

Real projects, real impact

Highlights from my professional experience.

Deloitte / Walmart

Enterprise Data Platform

Owning consumption layer in Spark-Scala, ingestion from 100+ SAP tables and Kafka brokers into GCS/BigQuery. Airflow DAGs with SLA monitoring and automated failure recovery.

Spark-Scala BigQuery Airflow Kafka GCP
ISI Kolkata / M.Tech Thesis

PCGT — Graph Transformer

Novel Graph Transformer replacing O(N²) global attention with partition-aware attention. Outperformed SGFormer on 10/12 benchmarks. Scales to 1.6M+ node graphs.

PyTorch PyG METIS Research
Personal Project

Notebook App

AI-powered note-taking app with rich editor, code blocks, math rendering, drawing canvas, PDF export. Integrates Gemini and Groq for AI writing. Deployed as PWA.

Next.js React Firebase AI TypeScript
Deloitte / Walmart

Pipeline Automation Tool

Python tool that standardized DAG and CCM configuration across new pipeline onboardings — column sensitivity detection, schema mapping, GCS provisioning.

Python Airflow Automation
Tata Power

IEX Energy Trading Forecast

ML forecasting model for day-ahead electricity trading on Indian Energy Exchange, achieving 88% accuracy with real-time PI Server integration.

Python ML FastAPI Power BI
Tata Power

Automated Reporting System

End-to-end automated Power BI dashboards for HR, customer operations, and engineering teams. REST APIs with FastAPI and Node.js serving 10+ dashboards.

Power BI FastAPI Node.js REST APIs
Allstate India

Payment Default Prediction

Optimized XGBoost model improving payment default prediction from 77% to 85% accuracy using geographic feature engineering and hyperparameter tuning.

XGBoost Python Feature Eng.

Let's connect

Open to Data Engineering opportunities at product companies.