CV

Author

Zhe Cui

Published

July 22, 2026

Summary

Applied scientist focused on designing measurement systems that align large-scale ML with human preferences — human-feedback data, LLM-as-judge, and the evaluation methodology that decides whether models are good enough to ship. I found and drive cross-functional initiatives that turn sparse, self-selected human judgments into bias-aware ground truth, and validate LLM judges against human ground truth across policy areas, languages, and cultures — spanning measurement validity, selection-bias correction, sparse-to-dense label modeling, and benchmark design, adjudication, and drift monitoring. Built on deep production-ML experience — retrieval, ranking, reranking, experimentation, causal inference — the components now central to RAG, tool selection, and judge ensembles.

Skills

Category Details
Evaluation LLM Evaluation & LLM-as-Judge, Human Preference & Feedback Data, Measurement Validity (construct / convergent / discriminant / incremental), Model–Human Agreement (Cohen’s/Fleiss’ kappa, precision / recall), Benchmark Design, Drift Monitoring
ML & Statistics Foundation Models, PyTorch, Causal Inference, Selection-Bias & Propensity Modeling, Empirical Bayes, Experimental Design, Reward/Preference Modeling
Domains Recommender Systems, Ranking & Retrieval, Multi-objective Optimization, Long-term Value Measurement, Multilingual & Cross-Cultural Evaluation
Programming Python, SQL, R

Industry Experience

Senior Applied Scientist, Snap (2024–present) | Seattle, WA

Individual contributor and technical lead. Originated and drove multiple multi-quarter, cross-functional evaluation and human-feedback initiatives from business case through production adoption — setting technical direction hands-on and aligning engineering, product, ranking, and UX partners through influence, not authority.

  • Founded & lead human-feedback measurement at scale: Built and shipped a production system converting sparse, self-selected human preference judgments into reliable ground truth for evaluating content-understanding and ranking models, framing the data-generating process (random assignment vs. self-selected response) to isolate where bias enters.
  • Owned the evaluation validity framework: Established construct, convergent, discriminant, and incremental validity — proving the signal measures the intended latent construct, not a gameable behavioral proxy — plus model–human disagreement diagnostics to guide deployment.
  • Directed bias characterization & correction: Reframed responder self-selection as a missing-label (not sampling) problem; modeled it with propensity models, verified covariate balance, and showed low- vs high-propensity labelers rank content consistently, so the signal generalizes beyond responders.
  • Architected sparse-to-dense label modeling: Designed a two-head multi-task model (selection × judgment) to extend sparse labels to every instance, with a hierarchical Empirical-Bayes estimator as an immediately usable bridge under label sparsity and cold-start.
  • Founded & lead LLM-as-judge evaluation for content safety: Built a reusable judge capability validating LLM decisions against human ground truth across policy areas and multiple languages, regions, and cultures — where “appropriate” is viewer- and locale-dependent, not a fixed rule. Anchored on dual benchmark sets (taxonomy coverage vs. production hard cases), blind multi-reviewer adjudication (Cohen’s/Fleiss’ kappa), and staged human-agreement → judge-alignment (precision/recall, FPR/FNR) → drift-monitoring metrics — delivered as a transferable playbook so new policies and locales are an iteration, not a new project.
  • Linked evaluation to model quality: Used the human-preference signal as an offline/online evaluation lens on production ranking — quantifying value tradeoffs behavioral metrics miss and surfacing content that engagement over-rewards.

Data Scientist, Meta (2017–2023) | Seattle, WA

Where I learned to build and evaluate large-scale ML systems — through experimentation, causal inference, and personalization.

  • Ranking, Retrieval & Personalization: Improved Facebook App Navigation and Groups ranking with two-tower neural retrieval models, evaluated through online experiments and implicit/explicit feedback — the offline/online evaluation and feedback-loop patterns now central to RAG, reranking, and judge ensembles.
  • Causal Inference for Query Latency: Applied observational causal inference (propensity-score modeling, regression discontinuity) to isolate the features most responsible for slow Ads API queries and prioritize fixes, reducing p99 latency ~10%.
  • Real-time Reliability Modeling: Built, deployed, and monitored GBDT models predicting API requests likely to OOM/timeout, routing expensive requests to an async tier — cutting daily advertiser API errors from 6–7% to 3%.
  • Failure-Pattern & Impact Analysis: Modeled Ads Manager session data to find error patterns predictive of uncompleted ad creation, and quantified ad-spend loss from campaign failures to prioritize incidents and regressions.

Data Scientist, Bell Canada (2015–2017) | Toronto, ON, Canada

  • Churn Modeling & Customer-Journey Analytics: Built ensemble churn models (Random Forest, GBDT, neural nets, GLMs) for targeted retention, and a Spark funnel/pattern-mining pipeline whose derived features lifted churn prediction by 20%.

Research Assistant, University of Toronto (2012–2015) | Toronto, ON, Canada

  • Resource Allocation in Backhaul-Constrained Small Cell Networks: Developed distributed algorithms in MATLAB and Python for cooperative resource allocation in 5G networks. Published and presented at CISS 2014 (IEE Explore, Paper, Slides).

Internships (2007–2012)

  • Apple (2011), Magnum Semiconductor (2011), ON Semiconductor (2008–2010): Projects in signal processing (audio, biomedical), machine learning, and embedded systems.

Education

  • Masters of Applied Science, Electrical Engineering, University of Toronto (2012–2015) Thesis: Resource Allocation in Backhaul Constrained Small-Cell Networks, Grade A
  • Bachelor of Applied Science, Electrical Engineering, University of Waterloo (2007–2012), Honors First Class
  • International Exchange, Electrical Engineering, National University of Singapore (2010)