- Built BlogTracker 2.0, a full-stack narrative intelligence platform analyzing 120K+ posts across 98 blogs, enabling real-time exploration of sentiment, topics, and influence.
- Analyzed millions of large-scale records using advanced SQL on PostgreSQL, identifying performance trends and optimization opportunities that informed data-driven recommendations.
- Designed NDJSON streaming pipelines and analytical data models that reduced ingestion-to-dashboard latency to sub-second repeat loads.
- Developed backend services with FastAPI to expose processed data and model outputs for downstream applications and dashboards.
- Designed pipeline workflows with task sequencing and dependency management patterns similar to Dagster / Airflow, improving pipeline reliability and maintainability.
- Integrated ML and AI pipelines (LLMs, clustering, embeddings) into data workflows, enabling advanced analytics and intelligent data products.
- Optimized large-scale processing using batching, indexing, and vectorized pipelines for scalable query and analysis performance.
- Implemented data validation, schema consistency enforcement, and reliable delivery mechanisms for downstream users.
Summary
Data Engineer and Scientist with 2+ years of experience building scalable ETL pipelines, statistical models, and machine learning-driven data products on large-scale datasets. Skilled in Python (PySpark, Pandas), advanced SQL, and modern data engineering practices including orchestration, data modeling, and cloud-based processing. Strong foundation in experimentation, causal inference, and BI — with proven ability to translate complex analyses into actionable business outcomes across distributed systems and cross-functional teams.
Work Experience
- Built and maintained ETL pipelines using Python and SQL to ingest, transform, and integrate data from Google Analytics, APIs, and structured sources.
- Designed and ran A/B tests and hypothesis-driven analyses to evaluate user engagement strategies.
- Developed statistical and machine learning models that improved campaign tracking accuracy by 40%.
- Built interactive dashboards in Tableau and Looker Studio, enabling self-service analytics on integrated datasets.
- Automated reporting workflows, reducing manual effort by 60% and enabling faster experimentation cycles.
- Partnered with product and UX teams to translate analytical findings into actionable product changes.
- Validated analytical datasets and SQL outputs through test-driven checks to ensure reliability across reporting pipelines.
- Built ML ranking and recommendation models (XGBoost, LightGBM, logistic regression, decision trees) to match healthcare workers to shifts, improving matching accuracy.
- Developed feature engineering pipelines transforming raw data into model-ready datasets incorporating experience, availability, and behavioral signals.
- Evaluated bias in ranking models (e.g., leave-based bias) and applied statistical techniques to mitigate fairness issues.
- Conducted causal analysis and hypothesis testing to validate workforce allocation strategies.
- Identified and corrected inefficient SQL queries, ensuring accurate data extraction and validation for downstream analytics.
- Delivered insights to stakeholders through SQL analysis and reporting, influencing operational decisions.
- Collected, cleaned, and analyzed datasets using Python and SQL for reporting and analysis tasks.
- Built dashboards and visualizations to communicate insights to technical and non-technical audiences.
- Improved data processing workflows by automating cleaning and transformation tasks.
- Performed exploratory data analysis to identify trends and patterns in datasets.