Deprecated: $wgMWOAuthSharedUserIDs=false is deprecated, set $wgMWOAuthSharedUserIDs=true, $wgMWOAuthSharedUserSource='local' instead [Called from MediaWiki\HookContainer\HookContainer::run in /var/www/html/w/includes/HookContainer/HookContainer.php at line 135] in /var/www/html/w/includes/Debug/MWDebug.php on line 372
ParisParl Corpus of Parliamentary Debates - MaRDI portal

Deprecated: Use of MediaWiki\Skin\SkinTemplate::injectLegacyMenusIntoPersonalTools was deprecated in Please make sure Skin option menus contains `user-menu` (and possibly `notifications`, `user-interface-preferences`, `user-page`) 1.46. [Called from MediaWiki\Skin\SkinTemplate::getPortletsTemplateData in /var/www/html/w/includes/Skin/SkinTemplate.php at line 677] in /var/www/html/w/includes/Debug/MWDebug.php on line 372

Deprecated: Use of QuickTemplate::(get/html/text/haveData) with parameter `personal_urls` was deprecated in MediaWiki Use content_navigation instead. [Called from MediaWiki\Skin\QuickTemplate::get in /var/www/html/w/includes/Skin/QuickTemplate.php at line 131] in /var/www/html/w/includes/Debug/MWDebug.php on line 372

ParisParl Corpus of Parliamentary Debates

From MaRDI portal



DOI10.5281/ZENODO.3819374Zenodo3819374MaRDI QIDQ6673594

Dataset published at Zenodo repository.

Author name not available (Why is that?)

Publication date: 10 May 2020

Copyright license: No records found.



The ParisParl Corpus of Parliamentary Debates, prepared in the PolMine Project, comprises all protocols of plenary sessions in the French Assemble nationale between 1996 and 2019. The corpus is built based on pdf documents issued by the Assemble nationale. The R package frappp has been used to extract structural information from the orginal text and to prepare an XML version of the corpus (preliminary TEI format). The structural annotation comprises speaker, party affiliation, parliamentary group affiliation, role, legislative period, session, date, interjections, year and agenda item. This release offers a linguistically annotated and indexed format of the corpus. As part of the corpus preparation pipeline, the data has been linguistically annotated (using the TreeTagger and StanfordNLP) and imported into the Corpus Workbench (CWB). The linguistic annotation comprises POS-tagging and lemmatization. This language resource is still very much in development and comes without any guarantees.






This page was built for dataset: ParisParl Corpus of Parliamentary Debates