OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
writing
- Benchmarks should shape the frontier, not just measure it · Apr 2026
- Closing the Evaluation Gap in Agentic AI · Feb 2026
- Design Principles for Iteratively Building AI Applications · Nov 2021
- Powerful Abstractions for Programmatically Building and Managing Training Sets (Stanford AI Lab blog) · Jun 2019
- Learning Math for Machine Learning · Aug 2018
- Building for the Blockchain (with Ramon Recuero) · Jan 2018
- How to Get into VR · May 2017
- How To Get Into Natural Language Processing · Jan 2017
publications
Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments
Agents’ Last Exam
Slice-based Learning: A Programming Model for Residual Learning in Critical Data Slices
Scene Graph Prediction with Limited Labels
Weakly supervised classification of rare aortic valve malformations using unlabeled cardiac MRI sequences
Full list on Google Scholar.
teaching
CS231N: Convolutional Neural Networks for Visual Recognition