Notice: Unexpected clearActionName after getActionName already called in /var/www/html/w/includes/context/RequestContext.php on line 333
PUlasso: High-Dimensional Variable Selection With Presence-Only Data - MaRDI portal

PUlasso: High-Dimensional Variable Selection With Presence-Only Data

From MaRDI portal
(Redirected from Publication:5978840)
Publication:3304856

DOI10.1080/01621459.2018.1546587zbMATH Open1437.62239arXiv1711.08129OpenAlexW2964352496WikidataQ91631047 ScholiaQ91631047MaRDI QIDQ3304856

Author name not available (Why is that?)

Publication date: 3 August 2020

Published in: (Search for Journal in Brave)

Abstract: In various real-world problems, we are presented with classification problems with positive and unlabeled data, referred to as presence-only responses. In this paper, we study variable selection in the context of presence only responses where the number of features or covariates p is large. The combination of presence-only responses and high dimensionality presents both statistical and computational challenges. In this paper, we develop the PUlasso algorithm for variable selection and classification with positive and unlabeled responses. Our algorithm involves using the majorization-minimization (MM) framework which is a generalization of the well-known expectation-maximization (EM) algorithm. In particular to make our algorithm scalable, we provide two computational speed-ups to the standard EM algorithm. We provide a theoretical guarantee where we first show that our algorithm converges to a stationary point, and then prove that any stationary point within a local neighborhood of the true parameter achieves the minimax optimal mean-squared error under both strict sparsity and group sparsity assumptions. We also demonstrate through simulations that our algorithm out-performs state-of-the-art algorithms in the moderate p settings in terms of classification performance. Finally, we demonstrate that our PUlasso algorithm performs well on a biochemistry example.


Full work available at URL: https://arxiv.org/abs/1711.08129



No records found.


No records found.








This page was built for publication: PUlasso: High-Dimensional Variable Selection With Presence-Only Data

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3304856)