Open to Data Engineer / Scientist roles

Hi, I'm Bhushan —
I build data systems that drive decisions.

Data Engineer / Data Scientist / Analytics Engineer

I design scalable ETL pipelines, ML-driven data products, and analytical platforms on large-scale datasets. Currently building BlogTracker 2.0 at COSMOS Lab (Univ. of Arkansas) — a narrative intelligence platform processing 120K+ posts with sentiment, topic, and influence modeling. Previously shipped experimentation pipelines, ranking models, and self-service dashboards that saved teams 60% of manual reporting time.

Little Rock, Arkansas, USA MS Applied Data Science, Indiana University 2+ years industry experience
120K+
Records processed in production ETL
60%
Manual reporting effort eliminated
40%
Campaign tracking accuracy improvement
3.8
Graduate GPA (MS Data Science, IU)
01 // About

Bridging engineering, statistics, and product impact.

I'm a data professional who works across the full stack — from designing PostgreSQL data models and PySpark ETL pipelines, to shipping production ML ranking systems with XGBoost and LightGBM, to building Tableau and Looker Studio dashboards that stakeholders actually use.

My MS in Applied Data Science (Statistics concentration) at Indiana University gave me the rigor for hypothesis testing, causal inference, and experimental design. My engineering background keeps me focused on what runs reliably in production, not just what works in a notebook.

I've worn analyst, scientist, and engineer hats across research labs and industry — and I enjoy the glue work: translating ambiguous business questions into data models, experiments, and measurable outcomes.

$ quick facts

  • Location Little Rock, AR
  • Education MS Data Science, IU
  • Focus Data Eng · ML · Analytics
  • Languages Python, SQL, R
  • Interests LLMs, experimentation, narrative intelligence
  • Status Open to full-time roles
02 // Tech Stack

Tools I reach for every day.

My stack spans data engineering, machine learning, and business intelligence — chosen to match the problem, not the hype cycle.

Languages

Python SQL R PySpark Bash

Data Engineering

ETL Pipelines Data Modeling Dagster Airflow FastAPI NDJSON Streaming

Machine Learning

XGBoost LightGBM Scikit-learn Clustering LDA Feature Engineering

Statistics & Experimentation

A/B Testing Hypothesis Testing Causal Inference Statistical Modeling Bias Evaluation

LLM & AI Systems

Embeddings LLM Evaluation Similarity Scoring Prompt Pipelines Topic Modeling

Data Systems

PostgreSQL BigQuery MySQL MS SQL Server Query Optimization

BI & Visualization

Tableau Power BI Looker Studio Excel (Advanced) Matplotlib

Cloud & Tooling

AWS (S3, EC2) Azure Data Factory Docker Git Linux
03 // Experience

Where I've built, measured, and shipped.

08/2025 – Present Current

Data Engineer I

COSMOS Lab · University of Arkansas
  • Architected BlogTracker 2.0, a full-stack narrative intelligence platform analyzing 120K+ posts across 98 blogs with real-time sentiment, topic, and influence exploration.
  • Designed NDJSON streaming pipelines and analytical data models reducing ingestion-to-dashboard latency to sub-second repeat loads.
  • Built FastAPI backend services exposing processed data and model outputs to downstream dashboards and applications.
  • Engineered Dagster / Airflow-style orchestration with task sequencing and dependency management, improving pipeline reliability and maintainability.
  • Integrated LLM, clustering, and embedding workflows into ETL pipelines — shipping intelligent data products, not just raw tables.
  • Optimized PostgreSQL analytics through indexing, batching, and vectorized processing for scalable query performance on large-scale workloads.
PythonPostgreSQLFastAPIPySparkLLMsEmbeddingsDagster
11/2023 – 05/2025

Data Scientist

Comet Lab · Indiana University
  • Built ETL pipelines (Python + SQL) integrating Google Analytics, APIs, and structured sources to power behavioral analysis.
  • Designed and ran A/B tests and hypothesis-driven analyses, improving campaign tracking accuracy by 40%.
  • Automated reporting workflows, reducing manual effort by 60% and enabling faster experimentation cycles.
  • Developed interactive Tableau and Looker Studio dashboards that translated insights into product and business decisions.
  • Partnered with product and UX teams to translate analytical findings into actionable changes and validated hypotheses.
PythonSQLTableauLooker StudioGoogle AnalyticsA/B Testing
05/2024 – 01/2025

Data Scientist

Favorite Healthcare Staffing
  • Built ML ranking models (XGBoost, LightGBM, logistic regression) to match healthcare workers to shifts, improving matching accuracy.
  • Designed feature engineering pipelines encoding experience, availability, and behavioral signals for model-ready datasets.
  • Evaluated and mitigated ranking bias (e.g., leave-based bias) using statistical fairness techniques.
  • Ran causal analysis and hypothesis testing to validate workforce allocation strategies and inform operational decisions.
  • Delivered SQL-driven insights and reports that influenced staffing and operational strategy.
XGBoostLightGBMScikit-learnSQLCausal InferenceFairness
07/2022 – 08/2022

Summer Intern

Ramrao Adik Institute of Technology · Mumbai, India
  • Collected, cleaned, and analyzed datasets using Python and SQL for internal reporting workflows.
  • Built dashboards and visualizations translating data into insights for technical and non-technical audiences.
  • Automated cleaning and transformation steps, improving the reliability of recurring analyses.
PythonSQLEDAReporting
05 // Projects

Selected personal & coursework projects.

A mix of ML systems, statistical analyses, and BI work — each solving a specific, well-scoped problem.

Fake News Detection
MACHINE LEARNING

Fake News Detection with ML

Classification pipeline that separates genuine news from fabricated stories using TF-IDF features and ensemble classifiers, aimed at combating digital misinformation.

PythonNLPScikit-learn
View
Diabetes Prediction
HEALTHCARE ML

Diabetes Prediction Model

Supervised learning model that identifies high-risk patients from clinical features — designed to support proactive interventions and personalized care.

PythonPandasML
View
Indian Election Analysis
STATISTICAL ANALYSIS

Indian Election Analysis 2019

Hypothesis-driven exploration of what actually drives a candidate's win chances beyond raw vote count — using data exploration, visualization, and statistical testing in R.

Rggplot2Hypothesis Testing
View
Amazing Mart Analysis
BI / TABLEAU

Amazing Mart Purchase Patterns

Interactive scatter-plot Tableau dashboard surfacing customer purchase patterns and segmentation signals to inform marketing strategy.

TableauEDA
View
UK Bank Analysis
BI / TABLEAU

UK Bank Usage by Region & Occupation

Tableau data story analyzing banking behavior segmented by geography and occupation — highlighting demographic trends that inform product design.

TableauStorytelling
View
FIFA 21 Analysis
EXPLORATORY ANALYSIS

FIFA 21 Player EDA

End-to-end exploratory analysis of the FIFA 21 player dataset using Python's data and visualization stack, surfacing trends for strategic insight.

PythonPandasSeaborn
View
Restaurant SQL Analysis
SQL / ANALYTICS

Restaurant Customer SQL Analysis

SQL-driven deep dive into restaurant transactional data, uncovering patron preferences, behavior patterns, and actionable service improvements.

SQLPostgreSQL
View
More on GitHub
GITHUB

More on GitHub →

Full project archive, notebooks, and coursework — from data engineering experiments to ML mini-projects and analytical write-ups.

@BhushanShelke3739
Visit
06 // Education

Formal training.

Indiana University, Indianapolis

M.S. Applied Data Science (Statistics Concentration)
08/2023 – 05/2025GPA 3.8 / 4.0

University of Mumbai

B.E. Electronics & Telecommunications Engineering
08/2019 – 05/2023GPA 3.7 / 4.0

Let's build something useful.

I'm actively looking for Data Engineer, Data Scientist, and Analytics Engineer roles. If you're hiring — or just want to chat about data systems, experimentation, or LLMs — reach out.