Bayesian Clustering via Fusing of Localized Densities

by   Alexander Dombowsky, et al.

Bayesian clustering typically relies on mixture models, with each component interpreted as a different cluster. After defining a prior for the component parameters and weights, Markov chain Monte Carlo (MCMC) algorithms are commonly used to produce samples from the posterior distribution of the component labels. The data are then clustered by minimizing the expectation of a clustering loss function that favours similarity to the component labels. Unfortunately, although these approaches are routinely implemented, clustering results are highly sensitive to kernel misspecification. For example, if Gaussian kernels are used but the true density of data within a cluster is even slightly non-Gaussian, then clusters will be broken into multiple Gaussian components. To address this problem, we develop Fusing of Localized Densities (FOLD), a novel clustering method that melds components together using the posterior of the kernels. FOLD has a fully Bayesian decision theoretic justification, naturally leads to uncertainty quantification, can be easily implemented as an add-on to MCMC algorithms for mixtures, and favours a small number of distinct clusters. We provide theoretical support for FOLD including clustering optimality under kernel misspecification. In simulated experiments and real data, FOLD outperforms competitors by minimizing the number of clusters while inferring meaningful group structure.


page 1

page 2

page 3

page 4


Kernel learning approaches for summarising and combining posterior similarity matrices

When using Markov chain Monte Carlo (MCMC) algorithms to perform inferen...

Bayesian clustering of high-dimensional data

In many applications, it is of interest to cluster subjects based on ver...

Anchored Bayesian Gaussian Mixture Models

Finite Gaussian mixtures are a flexible modeling tool for irregularly sh...

Estimating densities with nonlinear support using Fisher-Gaussian kernels

Current tools for multivariate density estimation struggle when the dens...

Distributed Bayesian clustering

In many modern applications, there is interest in analyzing enormous dat...

Forest Fire Clustering: Cluster-oriented Label Propagation Clustering and Monte Carlo Verification Inspired by Forest Fire Dynamics

Clustering methods group data points together and assign them group-leve...

Bayesian Clustering of Neural Activity with a Mixture of Dynamic Poisson Factor Analyzers

Modern neural recording techniques allow neuroscientists to observe the ...

Please sign up or login with your details

Forgot password? Click here to reset