EUDETECTOR: Leveraging Language Model to Identify EU-Related News. In: Companion Proceedings of the Web Conference 2021
Koustav Rudra
Danny Tran
Miroslav Shaltev
Abstract
News media reflects the present state of a country or region to its audiences. Media outlets of a region post different kinds of news for their local and global audiences. In this paper, we focus on Europe (precisely EU) and propose a method to identify news that has an impact on Europe from any aspect such as financial, business, crime, politics, etc. Predicting the location of the news is itself a challenging task. Most of the approaches restrict themselves towards named entities or handcrafted features. In this paper, we try to overcome that limitation i.e., instead of focusing only on the named entities (Europe location, politicians etc.) and some hand-crafted rules, we also explore the context of news articles with the help of pre-trained language model BERT. The auto-regressive language model based European news detector shows about 9-19% improvement in terms of F-score over baseline models. Interestingly, we observe that such models automatically capture named entities, their origin, etc; hence, no separate information is required. We also evaluate the role of such entities in the prediction and explore the tokens that BERT really looks at for deciding the news category. Entities such as person, location, organization turn out to be good rationale tokens for the prediction.
Rudra K., Tran D. and Shaltev M. (2021) EUDETECTOR: Leveraging Language Model to Identify EU-Related News. In: Leskovec J., Grobelnik M., Najork M.A., Tang J. & Zia L. (eds.). Companion Proceedings of the Web Conference 2021 (pp. 380-384). 10.1145/3442442.3452324
@inbook{Rudra2021,
Title = {EUDETECTOR: Leveraging Language Model to Identify EU-Related News},
Author = {Rudra, Koustav and Tran, Danny and Shaltev, Miroslav},
Editor = {Leskovec, J. and Grobelnik, M. and Najork, M.A. and Tang, J. and Zia, L.},
Year = {2021},
Pages = {380-384},
Doi = {10.1145/3442442.3452324},
Isbn = {978-1-4503-8313-4},
Booktitle = {Companion Proceedings of the Web Conference 2021},
Abstract = {News media reflects the present state of a country or region to its audiences. Media outlets of a region post different kinds of news for their local and global audiences. In this paper, we focus on Europe (precisely EU) and propose a method to identify news that has an impact on Europe from any aspect such as financial, business, crime, politics, etc. Predicting the location of the news is itself a challenging task. Most of the approaches restrict themselves towards named entities or handcrafted features. In this paper, we try to overcome that limitation i.e., instead of focusing only on the named entities (Europe location, politicians etc.) and some hand-crafted rules, we also explore the context of news articles with the help of pre-trained language model BERT. The auto-regressive language model based European news detector shows about 9-19% improvement in terms of F-score over baseline models. Interestingly, we observe that such models automatically capture named entities, their origin, etc; hence, no separate information is required. We also evaluate the role of such entities in the prediction and explore the tokens that BERT really looks at for deciding the news category. Entities such as person, location, organization turn out to be good rationale tokens for the prediction.},
}
TY - CHAP
AU - Rudra, Koustav
AU - Tran, Danny
AU - Shaltev, Miroslav
ED - Leskovec, Jure
ED - Grobelnik, Marko
ED - Najork, Marc A.
ED - Tang, Jie
ED - Zia, Leila
TI - EUDETECTOR: Leveraging Language Model to Identify EU-Related News
PY - 2021
SP - 380-384
DO - 10.1145/3442442.3452324
SN - 978-1-4503-8313-4
BT - Companion Proceedings of the Web Conference 2021
AB - News media reflects the present state of a country or region to its audiences. Media outlets of a region post different kinds of news for their local and global audiences. In this paper, we focus on Europe (precisely EU) and propose a method to identify news that has an impact on Europe from any aspect such as financial, business, crime, politics, etc. Predicting the location of the news is itself a challenging task. Most of the approaches restrict themselves towards named entities or handcrafted features. In this paper, we try to overcome that limitation i.e., instead of focusing only on the named entities (Europe location, politicians etc.) and some hand-crafted rules, we also explore the context of news articles with the help of pre-trained language model BERT. The auto-regressive language model based European news detector shows about 9-19% improvement in terms of F-score over baseline models. Interestingly, we observe that such models automatically capture named entities, their origin, etc; hence, no separate information is required. We also evaluate the role of such entities in the prediction and explore the tokens that BERT really looks at for deciding the news category. Entities such as person, location, organization turn out to be good rationale tokens for the prediction.
ER -
Details
Date03.06.2021
Pages380-384
Book TitleCompanion Proceedings of the Web Conference 2021
EditorsLeskovec J., Grobelnik M., Najork M.A., Tang J. and Zia L.