| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| en:cnk:dialekt [2022/01/05 15:54] – [DIALEKT corpus] martinawaclawicova | en:cnk:dialekt [2026/06/30 12:39] (current) – [Processing dialect recordings] michalkren |
|---|
| ===== Map of dialect regions in CR ===== | ===== Map of dialect regions in CR ===== |
| |
| {{:cnk:oblasti_ridsi_mod2.jpg?direct&500| Map of dialect regions in CR}} | {{:en:cnk:oblasti_ridsi_2021_wiki.png?direct&500| Map of dialect regions in CR}} |
| ====== Processing dialect recordings ====== | ====== Processing dialect recordings ====== |
| |
| Dialect material in the **DIALEKT** corpus is processed with two transcription tiers – dialectological and orthographic, see [[cnk:dialekt:pravidla|transcription principles]] (Czech only). The basic transcript is dialectological and is based on the rules for the transcription of scientific dialectological texts. The second transcription tier contains the orthographic transcription, which approaches the usual form of written texts and is comparable to the general rules established for spoken corpora in the Czech National Corpus (CNC). | Dialect material in the **DIALEKT** corpus is processed using the [[https://archive.mpi.nl/tla/elan|ELAN]] tool (developed in the Max Planck Institute for Psycholinguistics, Nijmegen((ELAN (Version 7.1) [Computer software]. (2026). Nijmegen: Max Planck Institute for Psycholinguistics. Retrieved from https://archive.mpi.nl/tla/elan |
| | ))) with two transcription tiers – dialectological and orthographic, see [[cnk:dialekt:pravidla|transcription principles]] (Czech only). The basic transcript is dialectological and is based on the rules for the transcription of scientific dialectological texts. The second transcription tier contains the orthographic transcription, which approaches the usual form of written texts and is comparable to the general rules established for spoken corpora in the Czech National Corpus (CNC). |
| **DIALEKT** is, similarly to the corpora **[[en:cnk:oral|ORAL]]** and **[[en:cnk:ortofon|ORTOFON]]** [[en:cnk:lemtag_mluv|lemmatized and morphologically tagged]]. Due to the extensive variability of dialect material and insufficient training data sets, the tagging and lemmatization process was extremely complicated, and it is necessary to keep this in mind when considering the outcome. | **DIALEKT** is, similarly to the corpora **[[en:cnk:oral|ORAL]]** and **[[en:cnk:ortofon|ORTOFON]]** [[en:cnk:lemtag_mluv|lemmatized and morphologically tagged]]. Due to the extensive variability of dialect material and insufficient training data sets, the tagging and lemmatization process was extremely complicated, and it is necessary to keep this in mind when considering the outcome. |
| |
| |
| Goláňová, H. – Waclawičová, M. – Komrsková, Z. – Lukeš, D. – Kopřivová, M. – Poukarová, P.: //DIALEKT: nářeční korpus, verze 1 z 2. 6. 2017//. Ústav Českého národního korpusu FF UK, Praha 2017. Retrieved from: http://www.korpus.cz\\ | Goláňová, H. – Waclawičová, M. – Komrsková, Z. – Lukeš, D. – Kopřivová, M. – Poukarová, P.: //DIALEKT: nářeční korpus, verze 1 z 2. 6. 2017//. Ústav Českého národního korpusu FF UK, Praha 2017. Retrieved from: http://www.korpus.cz\\ |
| | |
| | Goláňová, H. – Waclawičová, M. (2019): The DIALEKT corpus and its possibilities. Jazykovedný časopis, 70(2), 336-344. ISSN 0021-5597. |
| |
| Komrsková, Z. - Kopřivová, M. - Lukeš, D. - Poukarová, P. - Goláňová, H. (2017): New Spoken Corpora of Czech: ORTOFON and DIALEKT. //Jazykovedný časopis//, 68(2), 219-228. ISSN 0021-8897. | Komrsková, Z. - Kopřivová, M. - Lukeš, D. - Poukarová, P. - Goláňová, H. (2017): New Spoken Corpora of Czech: ORTOFON and DIALEKT. //Jazykovedný časopis//, 68(2), 219-228. ISSN 0021-8897. |
| Goláňová, H. (2015): A new dialect corpus: DIALEKT. In Katarína Gajdošová - Adriana Žáková (eds.): //Proceedings of the Eight International Conference Slovko 2015 (Natural Language Processing, Corpus Linguistics, Lexicography)//. Lüdenscheid: RAM-Verlag, 36-44. ISBN 978-3-942303-32-3.\\ | Goláňová, H. (2015): A new dialect corpus: DIALEKT. In Katarína Gajdošová - Adriana Žáková (eds.): //Proceedings of the Eight International Conference Slovko 2015 (Natural Language Processing, Corpus Linguistics, Lexicography)//. Lüdenscheid: RAM-Verlag, 36-44. ISBN 978-3-942303-32-3.\\ |
| |
| Goláňová, H. – Kopřivová, M. – Lukeš, D. – Štěpán, M. (2015): Kartografické a geografické zpracování dat z mluvených korpusů. In //Korpus – gramatika – axiologie//, 11, 42-54. ISSN: 1804-137X | |
| </WRAP> | </WRAP> |
| |
| Corpus compilation and project coordination was secured by //Hana Goláňová//, corpus preparation and proofreading of transcription by //Martina Waclawičová//, the orthographic transcription tier by //Zuzana Komrsková//, technical creation of the corpus by //David Lukeš// and lemmatization and morphological tagging was prepared by //Zuzana Komrsková//, //Marie Kopřivová//, //David Lukeš// and //Petra Poukarová//. | |
| |
| ===== Related links ===== | ===== Related links ===== |