Multilingual Part-of-Speech Tagging: Two Unsupervised Approaches

DOI10.1613/JAIR.2843zbMATH Open1192.68712arXiv1401.5695OpenAlexW2131134557WikidataQ129501916 ScholiaQ129501916MaRDI QIDQ3651489

Author name not available (Why is that?)

Publication date: 10 December 2009

Published in: (Search for Journal in Brave)

Abstract: We demonstrate the effectiveness of multilingual learning for unsupervised part-of-speech tagging. The central assumption of our work is that by combining cues from multiple languages, the structure of each becomes more apparent. We consider two ways of applying this intuition to the problem of unsupervised part-of-speech tagging: a model that directly merges tag structures for a pair of languages into a single sequence and a second model which instead incorporates multilingual context using latent variables. Both approaches are formulated as hierarchical Bayesian models, using Markov Chain Monte Carlo sampling techniques for inference. Our results demonstrate that by incorporating multilingual evidence we can achieve impressive performance gains across a range of scenarios. We also found that performance improves steadily as the number of available languages increases.

Full work available at URL: https://arxiv.org/abs/1401.5695

zbMATH Keywords

No records found.

Mathematics Subject Classification ID

No records found.

This page was built for publication: Multilingual Part-of-Speech Tagging: Two Unsupervised Approaches

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q3651489)