Date: 2026-07-21

The European project ATRIUM – Advancing Frontier Research in the Arts and Humanities aims to facilitate access to digital research infrastructures and support research in the arts and humanities across disciplines, languages and media. The project brings together four major European research infrastructures (DARIAH, ARIADNE, CLARIN and OPERAS) and addresses different stages of the research data lifecycle, from creation and processing to preservation, access and reuse.
As part of ATRIUM’s activities, and with the participation of members of the Digital Curation Unit at the Athena Research Centre, two new workflows have been developed and are now available through the SSH Open Marketplace. The workflows address two different but complementary areas of Digital Humanities research: the development of controlled vocabularies for historical map annotation and the geotagging of historical place names identified in texts.
The Controlled Vocabulary for Annotation of Historical Maps workflow provides a practical methodology for developing and maintaining controlled vocabularies that support historical map annotation. It guides research teams from defining the vocabulary’s purpose and scope to organising terms, writing definitions and usage rules, and examining existing external vocabularies in order to decide whether their terms should be reused or mapped to local terms. It also covers testing the vocabulary during annotation and preparing it for maintenance, publication and reuse.
The workflow addresses a key challenge that arises when annotation is carried out collaboratively: different annotators may use different terms for similar features or record the same information at different levels of detail. Controlled vocabularies help establish a more consistent and well documented annotation practice. Depending on the needs of each project, the resulting vocabulary may be used internally, published as a documented vocabulary package or released in a semantic format using standards such as SKOS (Simple Knowledge Organization System).
The second workflow, Geotagging of texts, focuses on identifying historical place names in texts and assigning geographic coordinates by linking them to reference gazetteers such as ToposText, Pleiades, GeoNames and the World Historical Gazetteer.
Resolving place references to persistent gazetteer URIs enables the retrieval of standardised geographic coordinates. The enriched data can then be integrated into other research applications and published as Linked Open Data (LOD).
The workflow combines automated and manual steps to improve geotagging accuracy. It also explores the use of Named Entity Recognition (NER) methods to identify place names, provides mechanisms for resolving ambiguous references and supports different input data formats.
Although they focus on different types of research material, the two workflows share a common objective: transforming complex content into consistent, well documented and reusable research data. They also highlight the role of controlled vocabularies, gazetteers and persistent identifiers in strengthening the interoperability and semantic interlinking of research information.