AplikaceAplikace
Nastavení

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
en:cnk:ortofon [2026/01/23 11:47] – [Morphological tagging of the ORTOFON corpus] krivanen:cnk:ortofon [2026/06/30 12:38] (current) michalkren
Line 1: Line 1:
 ====== Corpus of informal spoken Czech with multi-tier transcription: ORTOFON ====== ====== Corpus of informal spoken Czech with multi-tier transcription: ORTOFON ======
  
-The ORTOFON corpus captures spontaneous spoken language used in informal situations between speakers who know each other. It follows the [[en:cnk:oral|ORAL]] series of informal spoken Czech corpora in its data collection design. The recordings are transcribed in two tiers - orthographic and phonetic. Together with the [[en:cnk:dialekt|DIALEKT]] corpus, these are the first two spoken Czech corpora to have multi-tier transcription. Similar to the [[en:cnk:oral2013|ORAL2013]] corpus, speakers come from all over the Czech Republic and selected sociological information is collected about them. The corpus is lemmatized and morphologically tagged. The transcription is linked to the audio track and the audio can be played back in the KonText corpus interface. +The ORTOFON corpus captures spontaneous spoken language used in informal situations between speakers who know each other. It follows the [[en:cnk:oral|ORAL]] series of informal spoken Czech corpora in its data collection design. The recordings are transcribed in two tiers - orthographic and phonetic, using the [[https://archive.mpi.nl/tla/elan|ELAN]] tool, developed in the Max Planck Institute for Psycholinguistics, Nijmegen((ELAN (Version 7.1) [Computer software]. (2026). Nijmegen: Max Planck Institute for Psycholinguistics. Retrieved from https://archive.mpi.nl/tla/elan 
 +)). Together with the [[en:cnk:dialekt|DIALEKT]] corpus, these are the first two spoken Czech corpora to have multi-tier transcription. Similar to the [[en:cnk:oral2013|ORAL2013]] corpus, speakers come from all over the Czech Republic and selected sociological information is collected about them. The corpus is lemmatized and morphologically tagged. The transcription is linked to the audio track and the audio can be played back in the KonText corpus interface. 
  
 The ORTOFON corpus allows us to explore various aspects of spoken language, i.e. lexis, morphology, syntax, pragmatics, dialogue construction. The corpus is not primarily intended for dialectological ((The [[en:cnk:dialekt|DIALEKT]] corpus is intended for this kind of research.)) or phonetic research, even though a simplified phonetic transcription allows us to verify the existence of pronunciation or regional variants, or phenomena related to pronunciation. The ORTOFON corpus allows us to explore various aspects of spoken language, i.e. lexis, morphology, syntax, pragmatics, dialogue construction. The corpus is not primarily intended for dialectological ((The [[en:cnk:dialekt|DIALEKT]] corpus is intended for this kind of research.)) or phonetic research, even though a simplified phonetic transcription allows us to verify the existence of pronunciation or regional variants, or phenomena related to pronunciation.