AplikaceAplikace
Nastavení

Corpus SCRITTA

The corpus SCRITTA - School CoRpus of ITalian Texts by Adolescents - is a collection of texts in Italian written by teenagers at school. It was collected in 2020 and 2023 in Emilia-Romagna (Northern Italy) more precisely in schools of the cities of Bologna and Modena and their provinces. The corpus consists of 345 texts (126,382 tokens) written by pupils of the last year of junior high school (scuola secondaria di primo grado according to the Italian system). Each pupil contributed with one text.

The texts were written during ordinary school activities, as part of a written test, without the presence of the investigator.

The majority of the texts were handwritten and therefore transcribed manually; the remaining were collected in digital format. The whole corpus was read manually to remove any kind of personal information which could be used to recognise the author; anonimised elements can be traced by searching for the symbol @.

Metadata about the text authors' and their families were collected through a questionnaire and associated to the texts.

Text-related metadata include:

  • text type
  • writing modality (handwritten - typed)

School-related metadata:

  • area where the school is located (city centre - periphery - province)
  • average income of the area where the school is located
  • teacher id (alphanumeric code)

Author-related metadata include:

  • gender
  • age range
  • area of previous residence if relevant (other region of Italy; foreign country)
  • years spent in Emilia-Romagna
  • Italian region or country (if foreign) of birth of the parents
  • level of education of the parents
  • family language

The texts were automatically tokenised, lemmatised and annotated for PoS and dependency relations with italian-isdt-ud-2.17-251125. A manual revision of the annotation is planned: get in touch with the corpus developer if you would like to get involved!

Due to copyright-related reasons, the corpus cannot be downloaded and access to full texts cannot be granted to all users. Since different arrangements can be made, please get in touch with the corpus developer if you are interested.

For more details about the project, please see the works listed in the reference section.

Acknowledgements

This project was made possible thanks to the collaboration of many teachers, who believed in the dialogue between school and research, and hundreds of students.

Special thanks go to those who contributed to the transcription of the texts: Francesca Zucchini, Noemi Bellomo and Carlotta Sestito.

Contacts

Eleonora Zucchini, Masaryk University

eleonora.zucchini@mail.muni.cz

Studies based on this corpus

  • Zucchini, Eleonora. Italiano neostandard a scuola: Con un caso di studio sull'alternanza modale Bologna: Pàtron, 2025.
  • Grandi, Nicola and Eleonora Zucchini. Tratti neostandard nella scrittura formale giovanile. Un'indagine sulle scuole secondarie di Bologna. Rassegna Italiana di Linguistica Applicata. 2021, vol. 53(3), p. 121-258.
  • Perugini, Nicola and Eleonora Zucchini. Catene referenziali nel testo narrativo: un’indagine esplorativa sulla scrittura di alunni monolingui e plurilingui. Studi italiani di linguistica teorica e applicata. 2025, 54(1), p. 81-108.
  • Perugini, Nicola and Eleonora Zucchini. Reference in teenagers’ narratives between plurilingualism and restandardisation. In Ballarè, Silvia; Mauri, Caterina. CLUB Working Papers in Linguistics, Volume 9. Bologna: Circolo Linguistico dell'Università di Bologna, 2025, p. 99-123.
  • Zucchini, Eleonora. Elaborati scolastici a confronto: l’alternanza modale in testi di alunni italiani e ticinesi. Italiano a scuola. 2025, vol. 7, 39-60.