Controlling Complexity in Part-of-Speech Induction
From MaRDI portal
Publication:3093653
DOI10.1613/JAIR.3348zbMATH Open1234.68410arXiv1401.6131OpenAlexW3104758345WikidataQ129559913 ScholiaQ129559913MaRDI QIDQ3093653
Author name not available (Why is that?)
Publication date: 18 October 2011
Published in: (Search for Journal in Brave)
Abstract: We consider the problem of fully unsupervised learning of grammatical (part-of-speech) categories from unlabeled text. The standard maximum-likelihood hidden Markov model for this task performs poorly, because of its weak inductive bias and large model capacity. We address this problem by refining the model and modifying the learning objective to control its capacity via para- metric and non-parametric constraints. Our approach enforces word-category association sparsity, adds morphological and orthographic features, and eliminates hard-to-estimate parameters for rare words. We develop an efficient learning algorithm that is not much more computationally intensive than standard training. We also provide an open-source implementation of the algorithm. Our experiments on five diverse languages (Bulgarian, Danish, English, Portuguese, Spanish) achieve significant improvements compared with previous methods for the same task.
Full work available at URL: https://arxiv.org/abs/1401.6131
No records found.
No records found.
This page was built for publication: Controlling Complexity in Part-of-Speech Induction
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3093653)