Retail Media Pipelines
Large-scale data pipelines for retail transaction and media-performance datasets at Dunnhumby — built and optimized to be reliable at scale, not just functional.
- PB-scale datasets
- GCP native
- prod daily
available for opportunities / data engineer / 2026
I'm Sahil — a Google-certified Professional Data Engineer at Dunnhumby, building large-scale PySpark & SQL pipelines on GCP and BigQuery that turn raw retail data into decisions people can act on.
I'm a Junior Data Engineer at Dunnhumby, where I build and optimize large-scale data pipelines with PySpark and SQL for retail transaction and media-performance datasets — turning raw retail data into decisions people can act on.
Most of my work runs on Google Cloud Platform, especially BigQuery, processing datasets at scale and translating business requirements from cross-functional teams into pipelines that are reliable, not just functional. I hold Google's Professional Data Engineer certification, plus certs in Generative AI, LLMs, and Responsible AI.
Before data engineering I built ERP modules in Python at Odoo, and led the Alexa Developer Community at Chandigarh University — running workshops and hackathons that got other students building.
Daily driver for pipelines, scripting, and ML experiments.
Query optimization and transformation logic on large datasets.
Core tool for large-scale distributed data processing at work.
Google Certified — querying and managing datasets at scale.
Model building with scikit-learn and TensorFlow for applied problems.
Backend development for internal tools and side projects.
Large-scale data pipelines for retail transaction and media-performance datasets at Dunnhumby — built and optimized to be reliable at scale, not just functional.
Led a 3-person team to build a market-trend prediction model on historical price data, reaching 85% accuracy with classical ML.
Desktop app for encrypting and decrypting text messages — a hands-on exploration of practical cryptography and a clean Tkinter UI.
Steganography app that hides secret messages inside images through a simple interface — cryptography you can see, or rather, can't.
Build and optimize large-scale data processing pipelines using PySpark and SQL for retail transaction and media-performance datasets. Work extensively on GCP and BigQuery, collaborating with cross-functional teams to translate business requirements into scalable data engineering solutions.
Developed and customized Odoo ERP modules in Python for sales, inventory, and accounting. Used the Odoo ORM to design data models, implement business logic, and build forms, lists, and reports.
Organized developer workshops, hackathons, and seminars to strengthen the campus technology community, supporting engagement and technical learning among students.
Chandigarh University, Mohali · CGPA 7.57 / 10
Open to data engineering opportunities — or just a good conversation about pipelines, GCP, and generative AI.