inklap

Mixture‐Based Estimation of Multivariate Data Hypervolume

Luca Scrucca · Statistical Analysis and Data Mining: An ASA Data Science Journal · 2026

ABSTRACT Estimating the hypervolume occupied by multivariate data is a fundamental problem in statistics and data science, with applications ranging from ecology and machine learning to multi‐objective optimization and Bayesian inference. Traditional approaches rely on geometric approximations, kernel density estimation, or convex‐hull constructions, which often suffer from restrictive assumptions or do not scale well in higher dimensions. We introduce a novel methodology for hypervolume estimation based on finite Gaussian mixture models. The proposed approach defines the hypervolume as a high‐probability region of the fitted mixture density and estimates its volume using efficient Monte Carlo techniques, such as Latin hypercube sampling and importance sampling. An automatic, data‐driven procedure selects the density threshold that determines the region over which the hypervolume is computed. Across simulations, the proposed mixture‐based estimator proves broadly applicable and achieves accuracy, flexibility, and computational efficiency equal to or superior to those of existing methods. Applications to anomaly detection and ecological niche estimation illustrate

📖 افتح في inklap 🔗 DOI 📮 اطلب بحثاً