Bhushan Shelke

Data Engineer · Data Scientist · Analytics Engineer

Role-tailored resume variants available

The resume below combines my full experience. I also maintain targeted versions for Data Engineer, Data Scientist, and Data Analyst roles — available on request via bhushanshelke767@gmail.com.

Summary

Data Engineer and Scientist with 2+ years of experience building scalable ETL pipelines, statistical models, and machine learning-driven data products on large-scale datasets. Skilled in Python (PySpark, Pandas), advanced SQL, and modern data engineering practices including orchestration, data modeling, and cloud-based processing. Strong foundation in experimentation, causal inference, and BI — with proven ability to translate complex analyses into actionable business outcomes across distributed systems and cross-functional teams.

Work Experience

Data Engineer I
08/2025 – Present
COSMOS Lab — University of Arkansas, USA
Little Rock, AR
  • Built BlogTracker 2.0, a full-stack narrative intelligence platform analyzing 120K+ posts across 98 blogs, enabling real-time exploration of sentiment, topics, and influence.
  • Analyzed millions of large-scale records using advanced SQL on PostgreSQL, identifying performance trends and optimization opportunities that informed data-driven recommendations.
  • Designed NDJSON streaming pipelines and analytical data models that reduced ingestion-to-dashboard latency to sub-second repeat loads.
  • Developed backend services with FastAPI to expose processed data and model outputs for downstream applications and dashboards.
  • Designed pipeline workflows with task sequencing and dependency management patterns similar to Dagster / Airflow, improving pipeline reliability and maintainability.
  • Integrated ML and AI pipelines (LLMs, clustering, embeddings) into data workflows, enabling advanced analytics and intelligent data products.
  • Optimized large-scale processing using batching, indexing, and vectorized pipelines for scalable query and analysis performance.
  • Implemented data validation, schema consistency enforcement, and reliable delivery mechanisms for downstream users.
Data Scientist
11/2023 – 05/2025
Comet Lab — Indiana University, USA
Indianapolis, IN
  • Built and maintained ETL pipelines using Python and SQL to ingest, transform, and integrate data from Google Analytics, APIs, and structured sources.
  • Designed and ran A/B tests and hypothesis-driven analyses to evaluate user engagement strategies.
  • Developed statistical and machine learning models that improved campaign tracking accuracy by 40%.
  • Built interactive dashboards in Tableau and Looker Studio, enabling self-service analytics on integrated datasets.
  • Automated reporting workflows, reducing manual effort by 60% and enabling faster experimentation cycles.
  • Partnered with product and UX teams to translate analytical findings into actionable product changes.
  • Validated analytical datasets and SQL outputs through test-driven checks to ensure reliability across reporting pipelines.
Data Scientist
05/2024 – 01/2025
Favorite Healthcare Staffing, USA
Remote
  • Built ML ranking and recommendation models (XGBoost, LightGBM, logistic regression, decision trees) to match healthcare workers to shifts, improving matching accuracy.
  • Developed feature engineering pipelines transforming raw data into model-ready datasets incorporating experience, availability, and behavioral signals.
  • Evaluated bias in ranking models (e.g., leave-based bias) and applied statistical techniques to mitigate fairness issues.
  • Conducted causal analysis and hypothesis testing to validate workforce allocation strategies.
  • Identified and corrected inefficient SQL queries, ensuring accurate data extraction and validation for downstream analytics.
  • Delivered insights to stakeholders through SQL analysis and reporting, influencing operational decisions.
Summer Intern
07/2022 – 08/2022
Ramrao Adik Institute of Technology, India
Mumbai, India
  • Collected, cleaned, and analyzed datasets using Python and SQL for reporting and analysis tasks.
  • Built dashboards and visualizations to communicate insights to technical and non-technical audiences.
  • Improved data processing workflows by automating cleaning and transformation tasks.
  • Performed exploratory data analysis to identify trends and patterns in datasets.

Education

Indiana University, Indianapolis
08/2023 – 05/2025
Master of Science in Applied Data Science (Statistics Concentration)
GPA: 3.8 / 4.0
University of Mumbai, Mumbai
08/2019 – 05/2023
Bachelor of Engineering in Electronics & Telecommunications
GPA: 3.7 / 4.0

Skills & Tools

Programming
PythonSQL (Advanced)PySparkRPandasNumPyScikit-learnMatplotlib
Data Engineering
ETL PipelinesData ModelingDagsterAirflowFastAPINDJSON StreamingPipeline OptimizationData Validation
Machine Learning
XGBoostLightGBMLogistic RegressionDecision TreesClusteringLDAGradient BoostingFeature EngineeringModel Evaluation
Statistics
Hypothesis TestingA/B TestingCausal InferenceStatistical ModelingData Mining
LLM / AI
LLM EvaluationEmbeddingsSimilarity ScoringPrompt PipelinesTopic Modeling
Data Systems
PostgreSQLBigQueryMySQLMS SQL ServerSQL Optimization
Visualization / BI
TableauPower BILooker StudioExcel (Pivot Tables, Lookup)
Cloud & Tools
AWS (S3, EC2)Azure Data FactoryDockerGitSparkLinux