Data Engineer
Shreyas Shende
Data Engineer at Morgan Stanley working across the SDE, data, and applied-AI stack — production pipelines, agentic AI workflows, and internal tooling. Growing deeper into data science, with machine learning engineering as the long-term goal. MS in Computer Science from NJIT.

Beyond the Code
When I'm not building pipelines, this is where I'm usually at.
- Gaming
- Anime
- Hiking
- Swimming
Skills
A snapshot of the toolkit — the resume has the full list.
Data Engineering & Cloud
Programming & SQL
Analytics & BI
Machine Learning (Applied)
Collaboration & DevOps
Experience & Education
Hover a card (or tab to it) to read the details.
Data Engineer
Morgan Stanley · New York, NY · Dec 2025 – Present
- Designed and automated production ETL pipelines integrating data from 7+ enterprise APIs (Jira, Rally, OpenPages, etc.), processing ~35K records daily through scheduled AutoSys workflows, reducing pipeline failures by 30% and improving data reliability, resiliency, and quality through validation, structured logging, and fault-tolerant error handling.
- Built Agentic AI workflows leveraging enterprise GPT services to automate classification of 100+ monthly operational risk records — identifying cloud providers, cloud services, business impacts, and AI-related issues — with human-in-the-loop validation for audit readiness.
- Developed internal AI agents for Jira issue classification and automated pull request reviews, flagging security risks, coding standard violations, exposed secrets, and performance concerns before deployment.
- Automated CI/CD by configuring and maintaining 15 Jenkins pipelines and deployment workflows, while building operational Power BI dashboards and database fallback strategies over 2M+ records to improve monitoring and platform reliability.
Graduate Research Assistant (Master's Project)
New Jersey Institute of Technology · Newark, NJ · Jan 2025 – May 2025
- Built an end-to-end RNA-seq analysis pipeline in Python and PyTorch Geometric, automating data preprocessing, graph construction, model training, and feature selection across Cervical, Kidney, Alzheimer, and Lung cancer datasets; reduced feature space by >90% and improved classification accuracy by 5–10% over DESeq2 and EdgeR.
- Led a 3-member research team in collaboration with Brown University's Alpert Medical School to source and validate clinical RNA-seq datasets, resulting in a co-authored research paper demonstrating the pipeline's adaptability across diverse genomic datasets.
Platform Engineering & Automation Intern
Vendorpass (FIS Global) · Remote, US · May 2024 – Nov 2024
- Automated Python and Bash ETL workflows, reducing manual mapping effort by 90% and improving pipeline reliability.
- Built a Streamlit dashboard integrated with Oracle REST APIs to automate reporting across 20K+ records while implementing CI/CD validation and post-clone automation that reduced environment downtime from 4+ hours to 1 hour.
M.S. Computer Science
New Jersey Institute of Technology · Newark, NJ · Aug 2023 – May 2025
GPA 3.95/4.0
B.E. Computer Engineering
University of Pune · Pune, India · Aug 2019 – Jun 2023
GPA 3.66/4.0
Projects
A few things I've shipped recently.
Jul 2025 – Aug 2025
Containerized Python + Airflow pipeline to ingest, validate, and orchestrate raw ad campaign data into BigQuery. Databricks (PySpark) transformations clean and aggregate campaign KPIs (CTR, CPC, ROI), delivered through an interactive Power BI dashboard.
Pipeline
- Ad Campaign Data
- Airflow (Docker)
- BigQuery
- Databricks / PySpark
- Power BI
May 2025 – Jun 2025
Star-schema data model transforming nested JSON into Parquet for scalable analytics with Spark SQL. Interactive Power BI dashboard with DAX-based KPIs for engagement, churn, and revenue trends — cut reporting turnaround by 70%.
Pipeline
- Nested JSON
- Python ETL
- Parquet (Star Schema)
- Spark SQL
- Power BI (DAX)
Mar 2025 – Apr 2025
Scalable ETL pipeline using Airflow, Docker, and Celery to ingest Reddit data via APIs into Amazon S3, orchestrating transformations with AWS Glue and Athena. Automated ingestion of 100+ daily posts into Redshift for SQL-based trend and sentiment analytics.
Pipeline
- Reddit API
- Airflow + Celery
- Amazon S3
- AWS Glue / Athena
- Amazon Redshift
Certificates & Publications

AWS Certified Cloud Practitioner (CLF-C02)
Amazon Web Services

Microsoft Certified: Azure Data Fundamentals (DP-900)
Microsoft

Oracle Cloud Infrastructure 2025 Generative AI Certified Professional
Oracle

SnowPro Associate: Platform Certification
Snowflake
Publications
Google Scholar profile13
Citations
12 since 2021
2
h-index
2 since 2021
A Comprehensive Study on Simultaneous Localization and Mapping (SLAM): Types, Challenges, and Applications
A. Khole, A. Thakar, S. Shende, V. Karajkhede
2023 International Conference on Sustainable Computing and Smart Systems (ICSCSS), IEEE · 2023
A Compendium on Distributed Systems
A. Khole, A. Thakar, A. Kulkarni, H. Jadhav, S. Shende, V. Karajkhede
arXiv preprint arXiv:2302.03990 · 2023
RGE-GCN: Recursive Gene Elimination with Graph Convolutional Networks for RNA-seq based Early Cancer Detection
S. Shende, V. Narayanan, V. Fenn, Y. Huang, D. Goksuluk, G. Choudhary, et al.
arXiv preprint arXiv:2512.04333 · 2025
Real-Time Monocular SLAM: Accurate Localization and Mapping Using Point Map and Search-by-Projection Approach
A. Khole, A. Thakar, S. Shende, V. Karajkhede
Authorea Preprints · 2023
Let's build something
Open to Data Engineer roles and interesting problems. Reach out any time.
© 2026 Shreyas Shende. Built with Next.js.