| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| en:cnk:uvod [2026/05/19 11:09] – jankocek | en:cnk:uvod [2026/08/24 15:52] (current) – [Corpora of the Czech National Corpus project] michalkren |
|---|
| | **Parallel corpora** |||||| | | **Parallel corpora** |||||| |
| | [[en:cnk:intercorp|InterCorp]] ([[en:cnk:intercorp:verze16|release 16]], [[en:cnk:intercorp:verze16ud|release 16ud]]) | 5.3G | (✓) | (✓) | 2008–2024 | versioned parallel corpus for 61 languages | | | [[en:cnk:intercorp|InterCorp]] ([[en:cnk:intercorp:verze16|release 16]], [[en:cnk:intercorp:verze16ud|release 16ud]]) | 5.3G | (✓) | (✓) | 2008–2024 | versioned parallel corpus for 61 languages | |
| | [[en:cnk:romcro|RomCro 2.0]] | 19.4M | (✓) | (✓) | 2026 | parallel corpus of Romance languages and Croatian | | | [[en:cnk:romcro|RomCro 2.0]] | 19.4M | ✓ | ✓ | 2026 | parallel corpus of Romance languages and Croatian | |
| | [[en:cnk:psalm77|Psalm 77]] | 10k | (✓) | (✓) | 2023 | parallel corpus of 11 versions of Psalm 77 in Romanian, Church Slavonic and Greek | | | [[en:cnk:psalm77|Psalm 77]] | 10k | (✓) | (✓) | 2023 | parallel corpus of 11 versions of Psalm 77 in Romanian, Church Slavonic and Greek | |
| | **Comparable corpora** |||||| | | **Comparable corpora** |||||| |
| | [[en:cnk:nkjp|NKJP_1M]] | 1M | ✓ | ✓ | 2018 | manually annotated one-million subcorpus of the National Corpus of Polish | | | [[en:cnk:nkjp|NKJP_1M]] | 1M | ✓ | ✓ | 2018 | manually annotated one-million subcorpus of the National Corpus of Polish | |
| | [[en:cnk:obc|OBC]] | 24M | ✗ | ✓ | 2021 | [[http://fedora.clarin-d.uni-saarland.de/oldbailey/index.html|Old Bailey Corpus]], trial proceedings from 1720--1913 | | | [[en:cnk:obc|OBC]] | 24M | ✗ | ✓ | 2021 | [[http://fedora.clarin-d.uni-saarland.de/oldbailey/index.html|Old Bailey Corpus]], trial proceedings from 1720--1913 | |
| | | [[en:cnk:scritta|SCRITTA]] | 120k | ✓ | ✓ | 2026 | school corpus of Italian texts by adolescents | |
| ^ <fs large>Corpora generated by large language models (LLMs)</fs> ^^^^^^ | ^ <fs large>Corpora generated by large language models (LLMs)</fs> ^^^^^^ |
| ^ corpus ^ size (word count) ^ lemmas ^ morphological tags ^ year ^ characteristic features ^ | ^ corpus ^ size (word count) ^ lemmas ^ morphological tags ^ year ^ characteristic features ^ |