Evento
#Seminarios DaSCI

Aggregating Weak Annotations from Crowds

Ponente: Edwin Simpson, es profesor asociado en la Universidad de Bristol, donde trabaja en el procesamiento interactivo del lenguaje natural. Su investigación se centra en el aprendizaje a partir de datos escasos y poco fiables, incluida la retroalimentación de los usuarios, y adapta los enfoques bayesianos a temas como la argumentación, el resumen y el etiquetado de secuencias. Anteriormente, realizó un postdoctorado en la Universidad Técnica de Darmstadt, Alemania, y completó su doctorado en la Universidad de Oxford sobre los métodos bayesianos para agregar datos de origen colectivo.

Resumen: Current machine learning methods are data hungry. Crowdsourcing is a common solution to acquiring annotated data at large scale for a modest price. However, the quality of the annotations is highly variable and annotators do not always agree on the correct label for each data point. This talk presents techniques for aggregating crowdsourced annotations using preference learning and classifier combination to estimate gold-standard rankings and labels, which can be used as training data for ML models. We apply approximate Bayesian approaches to handle noise, small amounts of data per annotator, and provide a basis for active learning. While these techniques are applicable to any kind of data, we demonstrate their effectiveness for natural language processing tasks.

GrabaciónAggregating Weak Annotations from Crowds-Recordings (in English)

Aggregating Weak Annotations from CrowdsDescarga