European Multilingual News Articles Dataset with Topic Annotation (Q6709015)
From MaRDI portal
| This is the item page for this Wikibase entity, intended for internal use and editing purposes. Please use this page instead for the normal view: European Multilingual News Articles Dataset with Topic Annotation |
Dataset published at Zenodo repository.
| Language | Label | Description | Also known as |
|---|---|---|---|
| English | European Multilingual News Articles Dataset with Topic Annotation |
Dataset published at Zenodo repository. |
Statements
The European Multilingual News Articles Dataset is composed of over 18 million European news articles coming from 205 media outlets belonging to 27 European countries (i.e., all EU countries belonging to the European Union) with the addition of the United Kingdom. Articles range in a time period from 2017 to 2021 and are written in their original languages, for a total of 23 different languages included. After selecting reliable, nationwide European media outlets, each article (i.e., title, textual content, URL, and date and time of publication) was extracted from the Common Crawl News Corpus, which contains petabytes of raw web page data collected since 2016. The dataset is released without any text pre-processing other than a cleanup of XML tags. Further, we enriched it by adding several media metadata (e.g., frequency of publication, distribution area, language, type of media). Moreover, we enhanced the dataset by adding - whenever possible - article-level topic annotation by using articles' URLs as a proxy of the topic discussed. In the end, we were able to assign a topic to over 4 million articles (33 unique topics, e.g., politics, sport, entertainment), thus 23.2% of the entire dataset. Further, from URLs, we also extract the types of over 4 million articles (15 unique article types, e.g., news, international, multimedia).
0 references
17 December 2023
0 references
Version v1
0 references