AplikaceAplikace
Nastavení

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
en:cnk:orator [2025/06/06 13:34] – [Corpus composition and data acquisition] martinawaclawicovaen:cnk:orator [2026/06/30 12:39] (current) michalkren
Line 3: Line 3:
 The ORATOR corpus contains monologues by native Czech speakers. The typical situations include a lecture, instruction, guided tour, welcome address, sermon etc. The speech is usually prepared and the speaker has to fit within the given time frame. To our knowledge, there is no corpus with this kind of data available for Czech. The ORATOR corpus contains monologues by native Czech speakers. The typical situations include a lecture, instruction, guided tour, welcome address, sermon etc. The speech is usually prepared and the speaker has to fit within the given time frame. To our knowledge, there is no corpus with this kind of data available for Czech.
  
-Transcription rules, linking to the corresponding audio track and most metadata follow the [[en:cnk:ortofon|ORTOFON]] and [[en:cnk:oral|ORAL]] corpora, structural attributes used in ORATOR are described [[pojmy:atributy_strukturni|here]] (Czech only). The corpus is [[en:cnk:lemtag_mluv|lemmatized and morphologically tagged]] in the same way as the ORAL and ORTOFON corpora. The corpus is not balanced in any way.+Transcription rules, linking to the corresponding audio track (using the [[https://archive.mpi.nl/tla/elan|ELAN]] tool, developed in the Max Planck Institute for Psycholinguistics, Nijmegen((ELAN (Version 7.1) [Computer software]. (2026). Nijmegen: Max Planck Institute for Psycholinguistics. Retrieved from https://archive.mpi.nl/tla/elan 
 +)))  and most metadata follow the [[en:cnk:ortofon|ORTOFON]] and [[en:cnk:oral|ORAL]] corpora, structural attributes used in ORATOR are described [[pojmy:atributy_strukturni|here]] (Czech only). The corpus is [[en:cnk:lemtag_mluv|lemmatized and morphologically tagged]] in the same way as the ORAL and ORTOFON corpora. The corpus is not balanced in any way.
  
 <WRAP 45%> <WRAP 45%>
Line 28: Line 29:
 ===== Morphological tagging of the ORATOR corpus ===== ===== Morphological tagging of the ORATOR corpus =====
  
-The ORATOR v3 corpus is automatically [[en:pojmy:tag|annotated]] with [[en:cnk:syn2020#morphological_tagging|a new morphological tag]] according to the SYN2020 standard. It recognizes [[en:cnk:syn2020#multiple_lemmatization_and_tagging_aggregate|aggregates]] (e.g., //vidělas//, //zač//), uses [[en:cnk:syn2020|double-level lemmatization]], and has a verb tag ([[en:cnk:syn2020#verb_tagging_verbtag|verbtag]]). +The ORATOR v3 corpus is automatically [[en:pojmy:tag|annotated]] with [[en:cnk:syn2020#morphological_tagging|a new morphological tag]] according to the [[en:cnk:anotacni_standard_cnk|unified CNC annotation scheme]]. It recognizes [[en:cnk:syn2020#multiple_lemmatization_and_tagging_aggregate|aggregates]] (e.g., //vidělas//, //zač//), uses [[en:cnk:syn2020|double-level lemmatization]], and has a verb tag ([[en:cnk:syn2020#verb_tagging_verbtag|verbtag]]). 
  
-Substandard variants and forms typical of dialects and spontaneous speech are also tagged in the corpus. Special variants of words are distinguished by their own sublemma (e.g. //poslúchat// under the lemma //poslouchat//)special forms tagged only in the spoken corpus have the number 9 in the last tag position (e.g. the form //jezdijó// has the tag  ''%%VB-P---3P-AAI-9%%''). +Substandard variants and forms typical of dialects and spontaneous speech are also tagged in the corpus (according to the ORTOFON corpussee [[en:cnk:ortofon#morphological_tagging_of_the_ortofon_corpus|Morphological tagging of the ORTOFON corpus]]). 
  
 The following specific tags are used in the first tag position (word type): The following specific tags are used in the first tag position (word type):
Line 42: Line 43:
 Note: The anonymised sections are specified on a basic level ''%%word%%'': NP – surname, NJ – first name, NN – nickname, NM – place name, NO – other proper names, NT – last two digits of the telephone number. Note: The anonymised sections are specified on a basic level ''%%word%%'': NP – surname, NJ – first name, NN – nickname, NM – place name, NO – other proper names, NT – last two digits of the telephone number.
  
-The ORAL v1, ORTOFON v1 and ORTOFON v2 corpora are tagged with the prior morphological tagset used until 2020. Detailed information on the annotation of these previously published corpora can be found on a [[en:cnk:lemtag_mluv|separate page]].+The ORATOR v2 corpus is tagged with the prior morphological tagset used until 2020. Detailed information on the annotation of these previously published corpora can be found on a [[en:cnk:lemtag_mluv|separate page]].
  
 ====== ORATOR v1 (2019) ====== ====== ORATOR v1 (2019) ======
Line 54: Line 55:
 ====== ORATOR v3 (2025) ====== ====== ORATOR v3 (2025) ======
  
-The ORATOR corpus in its third version contains the same recordings and transcripts as the second version (i.e. over 1.5 million tokens)but they are newly annotated according to the SYN2020 standard. The genphone attribute is also newly included in the corpus, indicating the automatically generated phonetic form of a word. In addition, several transcription corrections have been made.+The ORATOR corpus in its third version contains the same recordings and transcripts as the second version (i.e. over 1.5 million tokens) but annotated according to the new Unified CNC Annotation Scheme using a language model trained also on spoken data. The ''genphone'' attribute is also newly included in the corpus, indicating the automatically generated phonetic form of a word. In addition, several transcription corrections have been made.
 ===== How to cite ===== ===== How to cite =====