Services
Courses, tutorials, workshops, and competitions.
Service
NeurIPS 2026 · Competition
The Predictive AI Evaluation Competition
Predicting AI responses on unseen benchmarks from model and item information.
COLM 2026 · Workshop
AI Measurement Science: Toward Rigorous AI Evaluation
October 9, 2026 · San Francisco, USA
AACL-IJCNLP 2026 · Tutorial
A Tutorial on Measurement Science for Language Model Evaluation
Measurement science for LLM evaluation: validity, predictive measurement, and benchmark design and incentives.
2025–2026 · Workshop Series
Language Models for Underserved Communities (LM4UC)
Workshop Organizer
Stanford · Community
Computing & Society
Interdisciplinary community examining how computing and AI affect society.
Teaching
Stanford · 2026
CS321M: AI Measurement Science
Frameworks for evaluating and understanding AI systems. Covers measurement as predictive modeling, measurement validity, and benchmark design and governance.
Teaching Fellow
Stanford · 2023–2026
CS329H: Machine Learning from Human Preferences
Learning and optimizing systems from human preferences, including choice models, reinforcement learning from human feedback, assistance, and fairness.
Teaching Fellow