Publication View

Active Learning for Part-of-Speech Tagging: Accelerating Corpus Annotation (2008)

Abstract
In the construction of a part-of-speech annotated corpus, we are constrained by a fixed budget. A fully annotated corpus is required, but we can afford to label only a subset. We train a Maximum Entropy Markov Model tagger from a labeled subset and automatically tag the remainder. This paper addresses the question of where to focus our manual tagging efforts in order to deliver an annotation of highest quality. In this context, we find that active learning is always helpful. We focus on Query by Uncertainty (QBU) and Query by Committee (QBC) and report on experiments with several baselines and new variations of QBC and QBU, inspired by weaknesses particular to their use in this application. Experiments on English prose and poetry test these approaches and evaluate their robustness. The results allow us to make recommendations for both types of text and raise questions that will lead to further inquiry. 1

Publication details
Download http://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.73.7658
Source http://james.jlcarroll.net/publications/LAW2007-Final.pdf
Contributors CiteSeerX
Repository CiteSeerX - Scientific Literature Digital Library and Search Engine (United States)
Type text
Language English
Relation 10.1.1.16.3103, 10.1.1.47.5102, 10.1.1.19.8901, 10.1.1.20.8521, 10.1.1.30.3233, 10.1.1.52.2415, 10.1.1.28.9963, 10.1.1.47.5135, 10.1.1.128.6885, 10.1.1.14.13, 10.1.1.137.6296, 10.1.1.111.7807