Bayesian Clustering via Fusing of Localized Densities

03/31/2023
by   Alexander Dombowsky, et al.
0

Bayesian clustering typically relies on mixture models, with each component interpreted as a different cluster. After defining a prior for the component parameters and weights, Markov chain Monte Carlo (MCMC) algorithms are commonly used to produce samples from the posterior distribution of the component labels. The data are then clustered by minimizing the expectation of a clustering loss function that favours similarity to the component labels. Unfortunately, although these approaches are routinely implemented, clustering results are highly sensitive to kernel misspecification. For example, if Gaussian kernels are used but the true density of data within a cluster is even slightly non-Gaussian, then clusters will be broken into multiple Gaussian components. To address this problem, we develop Fusing of Localized Densities (FOLD), a novel clustering method that melds components together using the posterior of the kernels. FOLD has a fully Bayesian decision theoretic justification, naturally leads to uncertainty quantification, can be easily implemented as an add-on to MCMC algorithms for mixtures, and favours a small number of distinct clusters. We provide theoretical support for FOLD including clustering optimality under kernel misspecification. In simulated experiments and real data, FOLD outperforms competitors by minimizing the number of clusters while inferring meaningful group structure.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
09/27/2020

Kernel learning approaches for summarising and combining posterior similarity matrices

When using Markov chain Monte Carlo (MCMC) algorithms to perform inferen...
research
06/04/2020

Bayesian clustering of high-dimensional data

In many applications, it is of interest to cluster subjects based on ver...
research
05/21/2018

Anchored Bayesian Gaussian Mixture Models

Finite Gaussian mixtures are a flexible modeling tool for irregularly sh...
research
07/12/2019

Estimating densities with nonlinear support using Fisher-Gaussian kernels

Current tools for multivariate density estimation struggle when the dens...
research
03/31/2020

Distributed Bayesian clustering

In many modern applications, there is interest in analyzing enormous dat...
research
03/22/2021

Forest Fire Clustering: Cluster-oriented Label Propagation Clustering and Monte Carlo Verification Inspired by Forest Fire Dynamics

Clustering methods group data points together and assign them group-leve...
research
05/21/2022

Bayesian Clustering of Neural Activity with a Mixture of Dynamic Poisson Factor Analyzers

Modern neural recording techniques allow neuroscientists to observe the ...

Please sign up or login with your details

Forgot password? Click here to reset