Revision 9 as of 2011-04-11 14:46:07

Clear message
Locked History Actions

PL196x

Polish language of the XX century sixties

This page is dedicated to the corpus of frequency dictionary of contemporary Polish. The original purpose of the corpus was to create a general frequency dictionary of contemporary Polish. The work started in 1967. Partial results were published between 1972 and 1977, the completed dictionary in 1990. The corpus was later augmented in various respects, both by manual editing and automated procedures.

Corpus data contain 10,000 samples divided into 5 parts: essays, news, scientific texts, fiction and plays. Every sample is approximately 50 words long, they all come from texts published between 1963 and 1967 and contain bibliographic description of its source. Each word is tagged with its base form and some morphological properties. Sentence boundaries are also marked.

In 2001 corpus authors agreed to publish the data in the Internet under GNU licence. This site presents corpus data in base and extended (enhanced) version as well as additional materials and corpus documentation.

Corpus documentation

Selected bibliography

  • Awramiuk Elżbieta. Wpływ odstępstw od segmentacji ortograficznej na wyniki statystyczne Słownika frekwencyjnego polszczyzny współczesnej. (In Polish). Roczniki Humanistyczne Uniwersytetu w Białymstoku 2001-2002, t. 49-50, z. 6, s. 31–43. Białystok 2001.

  • Bień, Janusz S.; Woliński, Marcin. Wzbogacony korpus Słownika frekwencyjnego polszczyzny współczesnej. (In Polish, EN: Enhanced corpus of the Frequency dictionary of contemporary Polish). [In:] Prace lingwistyczne dedykowane prof. Jadwidze Sambor. Jadwiga Linde-Usiekniewicz (ed.), pp. 6-10, Warszawa 2003, Faculty of Polish Philology, Warsaw University.

  • Kurcz, Ida; Lewicki, Andrzej; Sambor, Jadwiga; Woronczak, Jerzy. Vocabulary of contemporary Polish. Frequency lists. Volume I. Scientific texts. (In Polish). Warszawa, 1974. Warsaw University.

  • Kurcz, Ida; Lewicki, Andrzej; Sambor, Jadwiga; Woronczak, Jerzy. Vocabulary of contemporary Polish. Frequency lists. Volume II. News. (In Polish). Warszawa, 1974. Warsaw University.

  • Lewicki, Andrzej; Masłowski, Władysław; Sambor, Jadwiga; Woronczak, Jerzy. Vocabulary of contemporary Polish. Frequency lists. Volume III. Essays. (In Polish). Warszawa, 1975. Warsaw University.

  • Kurcz, Ida; Lewicki, Andrzej; Sambor, Jadwiga; Woronczak, Jerzy. Vocabulary of contemporary Polish. Frequency lists. Volume IV. Fiction. (In Polish). Warszawa, 1976. Warsaw University.

  • Kurcz, Ida; Lewicki, Andrzej; Sambor, Jadwiga; Woronczak, Jerzy. Vocabulary of contemporary Polish. Frequency lists. Volume V. Plays. (In Polish) Warszawa, 1977. Warsaw University.

  • Kurcz, Ida; Lewicki, Andrzej; Sambor, Jadwiga; Szafran, Krzysztof; Woronczak, Jerzy. Frequency dictionary of contemporary Polish. (In Polish). Kraków, 1990. Institute of Polish Philology, Polish Academy of Sciences.

  • Nazarczuk, Marta. Wstepne przygotowanie korpusu «Słownika frekwencyjnego polszczyzny współczesnej» do dystrybucji na CD-ROM. (In Polish, EN: Initial preparation of the corpus of Frequency dictionary of contemporary Polish for CD-ROM distribution). Master thesis prepared under supervision of prof. Janusz S. Bień. Warsaw, 1997. Institute of Polish Philology, Warsaw University. 59 pages, CD-ROM.

  • Ogrodniczuk, Maciej. Nowa edycja wzbogaconego korpusu słownika frekwencyjnego. (In Polish, EN: New edition of the Enhanced corpus of the Frequency dictionary). [In:] Językoznawstwo w Polsce. Stan i perspektywy. Stanisław Gajda (ed.) Institute of Polish Philology, Polish Academy of Sciences - Linguistics Committee, Opole University. Opole 2003, pp. 181-190. ISBN 83-86881-36-4.

  • Ogrodniczuk, Maciej. Rozszerzenie opisów morfologicznych w tekstach korpusu słownika frekwencyjnego polszczyzny współczesnej. (In Polish, EN: Augmenting the morphological description in the corpus of Frequency dictionary of contemporary Polish). [In:] Prace lingwistyczne dedykowane prof. Jadwidze Sambor. Jadwiga Linde-Usiekniewicz (ed.), pp. 164-168, Warszawa 2003, Faculty of Polish Philology, Warsaw University.

  • Ogrodniczuk, Maciej. Wykorzystanie SGML i TEI do zapisu polskich danych lingwistycznych. (In Polish, EN: Encoding of Polish linguistic data with SGML and TEI). Master thesis prepared under supervision of prof. Janusz S. Bień. Warsaw, 2000. Institute of Informatics, Warsaw University. 83 pages, CD-ROM.

  • Saloni, Zygmunt. Frequency dictionary of contemporary Polish. (In Polish). ComputerWorld, November 4th 1991, pp. 16-17.

  • Saloni, Zygmunt. Co skreślano i co dopisywano w korpusie Słownika frekwencyjnego polszczyzny współczesnej. (In Polish, EN: What was deleted and what was added in the frequency dictionary of contemporary Polish). pp. 381-391.

Corpus licence

Corpus data

Cluster

samples

"Raw"

Enhanced

TEI P4 XML

without codes

with codes

version

Style A: Scientific texts

1 MB

1,5 MB

1,1 MB

4,0 MB

10 MB

Style B: News

1 MB

1,5 MB

1,2 MB

3,9 MB

9 MB

Style C: Essays

1 MB

1,5 MB

1,2 MB

4,0 MB

10 MB

Style D: Fiction

1 MB

1,5 MB

1,1 MB

4,1 MB

11 MB

Style E: Plays

1 MB

1,5 MB

1,1 MB

4,4 MB

12 MB

Auxilliary files for the TEI P4-encoded XML version:

ISO image of the CD-ROM with most of the materials.

Concordances