EXMARaLDA Demo Corpus

The EXMARaLDA Demo Corpus contains examples of EXMARaLDA transcriptions in different languages. The EXMARaLDA Demo Corpus can be used to demonstrate and experiment with the functionality of the EXMARaLDA tools. Version 1.2 is explicitly designed to also demonstrate the use of the standard ISO 24624:2016 Language resource management — Transcription of spoken language.

The Demo Corpus is available in the following versions:

  • Version 1.0 was archived at the CLARIN-D repository of the Hamburg Centre for Language Corpora. It has been superseded by version 1.1 and is no longer available for download.
  • Version 1.1 is archived at the Zentrum für nachhaltiges Forschungsdatenmanagement of the University of Hamburg. It contains 26 communications and covers 13 languages.
  • Version 1.2 was created in the project Transcription+. It is a subset of version 1.1 (12 communications, 8 languages) with curated versions of all transcripts, including orthographic normalisation, lemmatisation and POS tagging in the ISO/TEI-Spoken transcript versions.

Published versions 1.1 and 1.2 can be downloaded from the ZFDM’s research data repository via https://doi.org/10.25592/uhhfdm.8363.
The most recent (unpublished) version is available from the GitHub repository at https://github.com/zumult-org/exmaraldademocorpus.
To use the corpus online, see AdWHH’s instance of the ZuMult corpus platform at https://transcription-plus.awhamburg.de/zumultapi/index.jsp