Size: 464
Comment:
|
Size: 2795
Comment:
|
Deletions are marked like this. | Additions are marked like this. |
Line 1: | Line 1: |
= The Polish Sejm Corpus = | = The Polish Sejm Corpus / Polski Korpus Sejmowy = |
Line 3: | Line 3: |
The Polish Sejm Corpus (PSC) is a large (currently 120M, target size: 300M segments) collection of stenographic transcripts of Polish Sejm sittings from 1-6 terms of office. | The Polish Sejm Corpus (PSC) is a large (300M-segment) collection of documents (stenographic transcripts, interpellations and questions) of [[http://www.sejm.gov.pl/english.html|Polish Sejm]] sittings from 1-8 terms of office. |
Line 5: | Line 5: |
The first edition of the PSC was prepared in October 2011 and contains automatically tokenized, lemmatized, morphosyntactically described and disambiguated transcripts of Sejm sessions saved in TEI P5 format. | The first edition of the PSC was prepared in October 2011 and was co-funded by [[CESAR]] European project. The current edition is being co-funded by [[CLARIN-PL-2|CLARIN-PL]]. |
Line 9: | Line 9: |
To be available soon. | The corpus contain transcripts of Sejm sessions saved in TEI P5 format. The resource contains automatically created annotation of: * utterance-level segmentation, * tokenization, * lemmatization, * disambiguated morphosyntactic description, * syntactic words, * syntactic groups, * named entities. === Whole corpus === * [[http://sejm.nlp.ipipan.waw.pl/static/PSC-all.tar|Interpellations and questions terms 3-8, sittings terms 1-8]], 25 GB === Divided by term and document type === ||<style="width:1%; align: center">Term ||<style="width:3%"> Years ||<style="width:10%"> Sittings ||<style="width:10%"> Interpellations and questions || ||<:> 1 ||<tablestyle="width:50%"> 1991–93 || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Posiedzenia-kad1.tar|1.0 GB]] || || ||<:> 2 || 1993–97 || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Posiedzenia-kad2.tar|2.6 GB]] || || ||<:> 3 || 1997–2001 || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Posiedzenia-kad3.tar|2.9 GB]] || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Interpelacje-kad3.tar|1.7 GB]] || ||<:> 4 || 2001–05 || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Posiedzenia-kad4.tar|3.4 GB]] || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Interpelacje-kad4.tar|2.4 GB]] || ||<:> 5 || 2005–07 || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Posiedzenia-kad5.tar|1.4 GB]] || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Interpelacje-kad5.tar|2.0 GB]] || ||<:> 6 || 2007–11 || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Posiedzenia-kad6.tar|3.2 GB]] || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Interpelacje-kad6.tar|4.8 GB]] || ||<:> 7 || 2011–15 || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Posiedzenia-kad7.tar|2.7 GB]] || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Interpelacje-kad7.tar|8.0 GB]] || ||<:> 8 || 2015– || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Posiedzenia-kad8.tar|0.7 GB]] || [[http://sejm.nlp.ipipan.waw.pl/static/PSC-Interpelacje-kad8.tar|1.1 GB]] || == Publications == <<BibMate(key,"ogro:12:lrec",omitYears=true)>> == Searching the corpus == Online search of the corpus is available at http://sejm.nlp.ipipan.waw.pl/. You can also use the Poliqarp image of the corpus: http://sejm.nlp.ipipan.waw.pl/static/PSC_poliqarp.tar.gz (826 MB). |
The Polish Sejm Corpus / Polski Korpus Sejmowy
The Polish Sejm Corpus (PSC) is a large (300M-segment) collection of documents (stenographic transcripts, interpellations and questions) of Polish Sejm sittings from 1-8 terms of office.
The first edition of the PSC was prepared in October 2011 and was co-funded by CESAR European project. The current edition is being co-funded by CLARIN-PL.
Corpus data
The corpus contain transcripts of Sejm sessions saved in TEI P5 format. The resource contains automatically created annotation of:
- utterance-level segmentation,
- tokenization,
- lemmatization,
- disambiguated morphosyntactic description,
- syntactic words,
- syntactic groups,
- named entities.
Whole corpus
Divided by term and document type
Term |
Years |
Sittings |
Interpellations and questions |
1 |
1991–93 |
|
|
2 |
1993–97 |
|
|
3 |
1997–2001 |
||
4 |
2001–05 |
||
5 |
2005–07 |
||
6 |
2007–11 |
||
7 |
2011–15 |
||
8 |
2015– |
Publications
Searching the corpus
Online search of the corpus is available at http://sejm.nlp.ipipan.waw.pl/. You can also use the Poliqarp image of the corpus: http://sejm.nlp.ipipan.waw.pl/static/PSC_poliqarp.tar.gz (826 MB).