Characterizing Out-of-Distribution Error via Optimal Transport

05/25/2023
by   Yuzhe Lu, et al.
0

Out-of-distribution (OOD) data poses serious challenges in deployed machine learning models, so methods of predicting a model's performance on OOD data without labels are important for machine learning safety. While a number of methods have been proposed by prior work, they often underestimate the actual error, sometimes by a large margin, which greatly impacts their applicability to real tasks. In this work, we identify pseudo-label shift, or the difference between the predicted and true OOD label distributions, as a key indicator to this underestimation. Based on this observation, we introduce a novel method for estimating model performance by leveraging optimal transport theory, Confidence Optimal Transport (COT), and show that it provably provides more robust error estimates in the presence of pseudo-label shift. Additionally, we introduce an empirically-motivated variant of COT, Confidence Optimal Transport with Thresholding (COTT), which applies thresholding to the individual transport costs and further improves the accuracy of COT's error estimates. We evaluate COT and COTT on a variety of standard benchmarks that induce various types of distribution shift – synthetic, novel subpopulation, and natural – and show that our approaches significantly outperform existing state-of-the-art methods with an up to 3x lower prediction error.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
02/10/2023

Predicting Out-of-Distribution Error with Confidence Optimal Transport

Out-of-distribution (OOD) data poses serious challenges in deployed mach...
research
08/04/2022

Interpretable Distribution Shift Detection using Optimal Transport

We propose a method to identify and characterize distribution shifts in ...
research
08/19/2020

Linearized Optimal Transport for Collider Events

We introduce an efficient framework for computing the distance between c...
research
10/04/2022

Tikhonov Regularization is Optimal Transport Robust under Martingale Constraints

Distributionally robust optimization has been shown to offer a principle...
research
06/07/2021

Measuring Generalization with Optimal Transport

Understanding the generalization of deep neural networks is one of the m...
research
03/18/2023

Uncertainty-Aware Optimal Transport for Semantically Coherent Out-of-Distribution Detection

Semantically coherent out-of-distribution (SCOOD) detection aims to disc...
research
10/04/2018

Generalizing the theory of cooperative inference

Cooperation information sharing is important to theories of human learni...

Please sign up or login with your details

Forgot password? Click here to reset