Portfolio

PRESENTATIONS

JupyterCon 2025 Speaker Badge

JupyterCon 2025 — Co-Presenter

“EDA Toolkit Tutorial on Streamlining Exploratory Data Analysis in Jupyter”
25-minute session co-presented with Leon Shpaner introducing EDA Toolkit, an open-source Python library for reproducible and efficient EDA in Jupyter notebooks.

  • Focus: Summary tables, automated profiling, distribution & crosstab visuals, export.

▶️ Watch the JupyterCon 2025 session on YouTube


Education Data Science Summit 2021 — Co-Presenter

“Using R and Machine Learning: How Did Having Internet and a Device at Home Impact Attendance in a Distance-Learning School Year?”
Applied analytics session (R + MySQL) modeling the relationship between home internet/device access and K–6 attendance during remote learning; included EDA, correlation, and linear regression.

  • Focus: Data-informed intervention insights for student attendance.
EDA Toolkit Logo

Co-Developer: EDA Toolkit (Open Source Python Library)

Co-developed EDA Toolkit, an open-source Python library for fast, reproducible exploratory data analysis.

  • Stack: Python, Pandas, statsmodels, SciPy
  • Key Features: EDA tools, contingency table creation, hypothesis testing, reproducible workflows
  • View Project →

A Journey in Volunteering with Survey Data Analysis

Collaborated with a nonprofit via Catchafire to analyze pre- and post-retreat leadership survey data, transforming raw responses into a structured format for meaningful insights. Designed efficient data staging and visualization workflows to summarize participant feedback and support program evaluation.

  • Stack: Python, Pandas, Matplotlib
  • Key Features:Data wrangling, survey response transformation, crosstab analysis, visual reporting
  • View Project →

Detecting Truncated Data Errors in SQL Server with Python: A Step-by-Step Guide

Proactively identified and resolved a SQL Server data truncation error by combining SQL Agent monitoring with Python-based data inspection. Used Pandas to quickly compare incoming data lengths against DDL specs, enabling a rapid fix before users were impacted.

  • Stack: SQL Server, Python, Pandas
  • Key Features:Proactive error detection, clipboard data analysis, DDL/data validation, rapid issue resolution
  • View Project →

From Traditional Pivot Tables to Crosstab in Python

Demonstrated an efficient workflow to replace manual Excel and Google Sheets pivot tables with automated, repeatable crosstab analysis in Python. Showcased the process using COVID-19 vaccine data from the California Open Data Portal.

  • Stack: Python, Pandas
  • Key Features:Automated pivot/crosstab analysis, CSV data ingestion, process reproducibility, streamlined data exploration
  • View Project →

University of San Diego Logo

Capstone MS, Applied Data Science (Jan 2021 – Dec 2022)

Developed machine learning models to predict English proficiency (ELPAC) levels for K–6 students using five years of California school district data, enabling early intervention for English Learners.

  • Stack: Python, scikit-learn, pandas, NumPy, Google Sheets, Excel
  • Key Features: Predictive modeling, feature engineering, actionable intervention insights
  • View Project →

How_Can_I_Help

How can I Help?

Let’s connect! If you need help with data science, analytics, automation, SQL, Python, or report development, I’m available for new projects and always open to new challenges—let’s talk!