We present an almost-linear time algorithm for the k-medoids problem that matches prior SOTA in clustering quality. Our solution has almost the same complexity as k-means and several advantages.
We show that selective classification, where models are allowed to abstain when they are uncertain, can fail to improve and even hurt accuracy over certain subpopulations of the data.
By tapping into knowledge stored explicitly in text corpora, retrieval helps tackle the inefficiency, opaqueness, and static nature of large language models.
How can we use machine learning to fix source code errors (e.g. in C, Python) for us? We introduce Break-It-Fix-It, a new unsupervised method to train code repair models.
A retrospective narrative from the Hazy research lab on our work in data-centric AI, and current efforts on engaging the broader machine learning community.
A novel computational tool for policymakers to assess the impacts of thousands of different mobility measures on predicted COVID-19 infections, helping them to navigate difficult tradeoffs between the economy and public health.