Data Analyst · Python · SQL · Tableau

Turning messy data into decisions.

I'm Yougi — an Informatics & Psychology student at UMass Amherst building dashboards, models, and data pipelines, including a live NYC collisions app over 27K records.

View work Get in touch Open to full-time · Summer 2027
Jul 2026 Automation / LLM Pipeline

Seed Miner

A scheduled GitHub Actions pipeline that samples one domain × technique pairing each weekday from an 89-cell matrix and has the Claude API draft a small, standalone analysis project for it. A manual promote/reject gate feeds back into sampling weights — decay-weighted, Laplace-smoothed promotion rates with a novelty bonus — so exploration never shuts off.

Seeds generated
40+
Matrix cells
89
Cadence
Weekdays
  • Python
  • Claude API
  • GitHub Actions
  • Weighted sampling
Jun 2026 Experimentation / Causal

When Average Effects Lie

Turns a 64,000-customer randomized email A/B test into a targeting strategy. The Women's email's +4.5pp average visit lift is really +7.3pp for prior women's-merch buyers and +1.1pp for everyone else. Balance checks, Lin-adjusted effects, FDR-corrected interactions, and uplift models yield a policy that roughly doubles net value while sending 41% fewer emails.

Customers
64,000
Experiment arms
3
Contacts saved
41%
  • Python
  • A/B testing
  • Causal inference
  • Uplift modeling
  • statsmodels
Jun 2026 Product Analytics

Spotify Listening Analysis

"Wrapped, but honest" — a SQL-first DuckDB pipeline over Spotify's Extended Streaming History: 30-minute sessionization, skip-rate and taste-concentration (HHI) metrics, artist cohort retention, and a two-proportion z-test on whether shuffle is skippier. Built-in QA reconciles every raw row, and it runs end-to-end on a synthetic sample so no personal data is committed.

SQL stages
5
Plays modeled
7,978
QA checks
4
  • SQL
  • DuckDB
  • Python
  • Cohort retention
  • Hypothesis testing
Jun 2026 Simulation / Full-Stack

LineLab — Decision Coach

A beginner-focused Teamfight Tactics coach built on an original Monte Carlo expected-value engine. Describe a game state and the FastAPI solver simulates thousands of rollouts, then collapses the EV table into one move — roll, level, save, or stabilize — with a plain-English why. Includes a playable Arena that grades every decision post-game.

Lessons
14
Modules
5
Coach verbs
4
  • Python
  • FastAPI
  • Monte Carlo
  • Next.js
  • TypeScript
  • Supabase
Dec 2025 Machine Learning

Bank Customer Churn Prediction

An end-to-end R report flagging at-risk customers across 10,000 records and 14 features. Logistic regression vs. Random Forest under 5-fold cross-validation, handling a 20.4% churn class imbalance, with decision thresholds tied to retention strategy.

Records
10,000
Features
14
Churn rate
20.4%
  • R
  • Classification
  • Random Forest
  • Class imbalance
  • Cross-validation
Aug 2025 Data Architecture

Plug — Campus Marketplace

Selected for Purdue's Market Readiness Incubator. As project manager, I designed the analytics-ready PostgreSQL (Supabase) backend — a relational schema with SQL CRUD plus multi-field search, sorting, and pagination, backed by integrity checks.

Backend
PostgreSQL
Queries
CRUD + search
Selected
Purdue Incubator
  • PostgreSQL
  • Supabase
  • SQL
  • Data Modeling
More projects on GitHub ↗
Focus
Analytics, ML, data engineering
Education
UMass Amherst — Informatics & Psychology
Graduating
May 2027
Based in
Amherst, MA
Projects shipped
8
Records analyzed
101K+
Survey participants
700+
Live apps
1
  1. Summer 2026 — Present

    Analysis Working Group Member — AI/ML and Brain

    NASA Open Science Data Repository · Remote

    Volunteer member of two of NASA OSDR's open-science Analysis Working Groups: AI/ML, which builds machine learning tools for spaceflight biology, and Brain. In Brain's Behavior & Neuroinflammation subgroup I work on linking a rodent behavioral phenotype from Rodent Research-1 to inflammasome-related neuroinflammation under spaceflight and ground-based radiation — pairing OSDR omics and behavioral datasets with environmental telemetry and RadLab radiation records. The subgroup is defining its research question and analysis plan.

  2. May 2025

    AI & Psychology Research Intern

    Spiritual Data · Remote

    Built a benchmarking harness comparing a proprietary model against baselines on 3+ metrics, designed evaluation splits and error slices to surface failure modes, and produced a stakeholder evaluation report aligned to business and compliance needs.

  3. Jun 2022

    Research Intern — Impact of Social Media

    UC Berkeley · Lawrence Hall of Science · Berkeley, CA

    Cleaned and prepared survey data from 700+ participants into analysis-ready datasets, then ran quantitative analysis (EDA, regression) on how social media use relates to public attitudes.

University of Massachusetts Amherst

Amherst, MA · Expected May 2027

B.S. Informatics & Psychology — Double Major

A double major bridging data and human behavior — informatics for the technical foundation (programming, databases, analysis) and psychology for the research methods and statistics behind asking sharper questions of the numbers.

Languages

  • Python
  • SQL
  • R
  • TypeScript
  • Java

Data

  • Pandas
  • NumPy
  • ETL
  • DuckDB
  • SQLite
  • PostgreSQL

Analytics / BI

  • Streamlit
  • Tableau
  • Power BI
  • Excel
  • KPI design

ML / Stats

  • scikit-learn
  • Classification
  • Regression
  • A/B testing
  • Causal inference
  • Uplift modeling
  • Monte Carlo
  • Model evaluation
  • Statistics

Tools

  • Git / GitHub
  • GitHub Actions
  • pytest
  • CI/CD
  • FastAPI
  • Caching
  • LangChain
  • Claude API

Let's build something with data.

Open to full-time data analyst roles and collaborations.