Skip to content
View CocoChengtw's full-sized avatar

Highlights

  • Pro

Block or report CocoChengtw

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
CocoChengtw/README.md

Hi, I'm Coco (HsiuWen Cheng)

Data Scientist / ML Engineer · 7 years building ML for Trust & Safety, fraud, and health data

I take ML from problem framing to production: defining the right metric, building models that handle messy, imbalanced real-world data, and measuring whether they actually changed outcomes.

  • Trust & Safety / Fraud — Senior DS at Trend Micro (scam SMS classification, deepfake detection); Data Scientist Intern at TikTok Integrity & Safety
  • Experimentation & causal inference — Scaled experimentation and campaign measurement for 200+ stakeholders at Far EasTone (~$430K/year value)
  • Health AI: time series & offline RL — CGM glucose time-series prediction; researcher at UCLA applying offline reinforcement learning to personalized CGM wear scheduling
  • MSBA, UCLA Anderson (2026)

Selected work

Project What it shows
Smishing Scam-Type Classifier Multilingual scam-type classification on 34k public smishing messages; template-aware splits expose a 3-point leakage gap, and confidence routing auto-labels 60% of traffic at 96.8% accuracy
Multi-Agent Dispatch QA (team) 4-person UCLA MSBA project extending a LangGraph demo into an audited multi-agent pipeline for specialty-medicine logistics. I built the deterministic AuditAgent (7 rules; failed plans loop back to the planner), the what-if ScenarioAgent, audit routing in the graph, a Gemini backend with an LLM call-budget guard, and tests
Food-101 CV Benchmark Systematic comparison of 14 architectures (ResNet → ConvNeXt → ViT); ConvNeXt-Base reached 87.9% top-1 with full fine-tuning
Deepfake Detection CNN forgery detection on public data, with threshold tuning from ROC/PR analysis
Lip-Sync Deepfake Detection LLM-guided video preprocessing; documents each iteration from 23% to 58% F1
Airbnb Market Intelligence Pipeline Bronze→Silver→Gold pipeline across 4 cities: Spark + Sedona geospatial joins, Airflow, Snowflake

Impact highlights from industry work (code proprietary)

  • Multi-layer transformer classifier (EN/JA) sorting scam SMS into 19 sub-categories under severe class imbalance, deployed for daily inference
  • Deepfake-detection false-positive fixes that cut user complaints from thousands to dozens

Toolkit

ML: PyTorch · Hugging Face · scikit-learn · time-series forecasting (LSTM) · offline RL (CQL, FQE)
Data & MLOps: Spark · Airflow · Snowflake · Databricks · FastAPI · Docker
Analysis: Causal inference (DiD) · A/B testing · Tableau · Looker Studio
Languages: English · 繁體中文

Contact: Website · LinkedIn · chenghsiuwen.tw@gmail.com

Pinned Loading

  1. smishing-scam-type-classifier smishing-scam-type-classifier Public

    Multilingual scam-type classification for SMS phishing, with template-aware evaluation, cluster-bootstrap CIs and confidence-based review routing.

    Python 1

  2. ucla-msba-s3/seewees-ai-agents-s3 ucla-msba-s3/seewees-ai-agents-s3 Public

    Python 1

  3. food101-cv-models food101-cv-models Public

    Systematic fine-tuning benchmark of 14 CNN and Vision Transformer architectures on Food-101, comparing full vs. partial fine-tuning (PyTorch).

    Python 2

  4. DeepfakeDetection DeepfakeDetection Public

    CNN-based video forgery (deepfake) detection on public data, with embedding analysis and threshold tuning from ROC/PR curves.

    Jupyter Notebook 1

  5. LipsyncDetection LipsyncDetection Public

    Lip-sync deepfake detection built on LipForensics, with LLM (Gemini) guided video segment selection and documented preprocessing experiments.

    Jupyter Notebook 1

  6. airbnb-market-intelligence-pipeline airbnb-market-intelligence-pipeline Public

    Bronze-Silver-Gold pipeline turning raw Airbnb listings into neighborhood market indicators across 4 US cities: Spark, Sedona geospatial joins, Airflow, Snowflake.

    Python 1