AplikaceAplikace
Nastavení

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
en:cnk:ortofon [2025/04/15 10:43] – [How to cite] michalkrenen:cnk:ortofon [2026/06/30 12:38] (current) michalkren
Line 1: Line 1:
 ====== Corpus of informal spoken Czech with multi-tier transcription: ORTOFON ====== ====== Corpus of informal spoken Czech with multi-tier transcription: ORTOFON ======
  
-The ORTOFON corpus captures spontaneous spoken language used in informal situations between speakers who know each other. It follows the [[en:cnk:oral|ORAL]] series of informal spoken Czech corpora in its data collection design. The recordings are transcribed in two tiers - orthographic and phonetic. Together with the [[en:cnk:dialekt|DIALEKT]] corpus, these are the first two spoken Czech corpora to have multi-tier transcription. Similar to the [[en:cnk:oral2013|ORAL2013]] corpus, speakers come from all over the Czech Republic and selected sociological information is collected about them. The corpus is lemmatized and morphologically tagged. The transcription is linked to the audio track and the audio can be played back in the KonText corpus interface. +The ORTOFON corpus captures spontaneous spoken language used in informal situations between speakers who know each other. It follows the [[en:cnk:oral|ORAL]] series of informal spoken Czech corpora in its data collection design. The recordings are transcribed in two tiers - orthographic and phonetic, using the [[https://archive.mpi.nl/tla/elan|ELAN]] tool, developed in the Max Planck Institute for Psycholinguistics, Nijmegen((ELAN (Version 7.1) [Computer software]. (2026). Nijmegen: Max Planck Institute for Psycholinguistics. Retrieved from https://archive.mpi.nl/tla/elan 
 +)). Together with the [[en:cnk:dialekt|DIALEKT]] corpus, these are the first two spoken Czech corpora to have multi-tier transcription. Similar to the [[en:cnk:oral2013|ORAL2013]] corpus, speakers come from all over the Czech Republic and selected sociological information is collected about them. The corpus is lemmatized and morphologically tagged. The transcription is linked to the audio track and the audio can be played back in the KonText corpus interface. 
  
 The ORTOFON corpus allows us to explore various aspects of spoken language, i.e. lexis, morphology, syntax, pragmatics, dialogue construction. The corpus is not primarily intended for dialectological ((The [[en:cnk:dialekt|DIALEKT]] corpus is intended for this kind of research.)) or phonetic research, even though a simplified phonetic transcription allows us to verify the existence of pronunciation or regional variants, or phenomena related to pronunciation. The ORTOFON corpus allows us to explore various aspects of spoken language, i.e. lexis, morphology, syntax, pragmatics, dialogue construction. The corpus is not primarily intended for dialectological ((The [[en:cnk:dialekt|DIALEKT]] corpus is intended for this kind of research.)) or phonetic research, even though a simplified phonetic transcription allows us to verify the existence of pronunciation or regional variants, or phenomena related to pronunciation.
Line 34: Line 35:
 ===== Morphological tagging of the ORTOFON corpus ===== ===== Morphological tagging of the ORTOFON corpus =====
  
-The ORTOFON v3 corpus is automatically [[en:pojmy:tag|annotated]] with [[en:cnk:syn2020#morphological_tagging|a new morphological tag]] according to the SYN2020 standard. It recognizes [[en:cnk:syn2020#multiple_lemmatization_and_tagging_aggregate|aggregates]] (e.g., //vidělas//, //zač//), uses [[en:cnk:syn2020|double-level lemmatization]], and has a verb tag ([[en:cnk:syn2020#verb_tagging_verbtag|verbtag]]). +The ORTOFON v3 corpus is automatically [[en:pojmy:tag|annotated]] with [[en:cnk:syn2020#morphological_tagging|a new morphological tag]] according to the [[en:cnk:anotacni_standard_cnk|unified CNC annotation scheme]]. It recognizes [[en:cnk:syn2020#multiple_lemmatization_and_tagging_aggregate|aggregates]] (e.g., //vidělas//, //zač//), uses [[en:cnk:syn2020|double-level lemmatization]], and has a verb tag ([[en:cnk:syn2020#verb_tagging_verbtag|verbtag]]). 
  
 Substandard variants and forms typical of dialects and spontaneous speech are also tagged in the corpus. Special variants of words are distinguished by their own sublemma (e.g. //poslúchat// under the lemma //poslouchat//), special forms tagged only in the spoken corpus have the number 9 in the last tag position (e.g. the form //jezdijó// has the tag  ''%%VB-P---3P-AAI-9%%'').  Substandard variants and forms typical of dialects and spontaneous speech are also tagged in the corpus. Special variants of words are distinguished by their own sublemma (e.g. //poslúchat// under the lemma //poslouchat//), special forms tagged only in the spoken corpus have the number 9 in the last tag position (e.g. the form //jezdijó// has the tag  ''%%VB-P---3P-AAI-9%%'').