Sandwich Boosting for Accurate Estimation in Partially Linear Models for Grouped Data

07/21/2023
by   Elliot H. Young, et al.
0

We study partially linear models in settings where observations are arranged in independent groups but may exhibit within-group dependence. Existing approaches estimate linear model parameters through weighted least squares, with optimal weights (given by the inverse covariance of the response, conditional on the covariates) typically estimated by maximising a (restricted) likelihood from random effects modelling or by using generalised estimating equations. We introduce a new 'sandwich loss' whose population minimiser coincides with the weights of these approaches when the parametric forms for the conditional covariance are well-specified, but can yield arbitrarily large improvements in linear parameter estimation accuracy when they are not. Under relatively mild conditions, our estimated coefficients are asymptotically Gaussian and enjoy minimal variance among estimators with weights restricted to a given class of functions, when user-chosen regression methods are used to estimate nuisance functions. We further expand the class of functional forms for the weights that may be fitted beyond parametric models by leveraging the flexibility of modern machine learning methods within a new gradient boosting scheme for minimising the sandwich loss. We demonstrate the effectiveness of both the sandwich loss and what we call 'sandwich boosting' in a variety of settings with simulated and real-world data.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
01/26/2019

On doubly robust estimation for logistic partially linear models

Consider a logistic partially linear model, in which the logit of the me...
research
08/17/2023

Average partial effect estimation using double machine learning

Single-parameter summaries of variable effects are desirable for ease of...
research
08/18/2017

A debiased distributed estimation for sparse partially linear models in diverging dimensions

We consider a distributed estimation of the double-penalized least squar...
research
03/12/2018

Partially Linear Spatial Probit Models

A partially linear probit model for spatially dependent data is consider...
research
05/19/2021

Latent Gaussian Model Boosting

Latent Gaussian models and boosting are widely used techniques in statis...
research
09/08/2021

Estimation for recurrent events through conditional estimating equations

We present new estimators for the statistical analysis of the dependence...
research
02/20/2022

KLLR: A scale-dependent, multivariate model class for regression analysis

The underlying physics of astronomical systems governs the relation betw...

Please sign up or login with your details

Forgot password? Click here to reset