1✉ Department of Biology, University of Graz, Universitaetsplatz 2, 8010 Graz, Austria.
2Department of Biology, University of Graz, Universitaetsplatz 2, 8010 Graz, Austria.
3Department of Biology, University of Graz, Universitaetsplatz 2, 8010 Graz, Austria.
4Department of Biology, University of Graz, Universitaetsplatz 2, 8010 Graz, Austria.
5Department of Biology, University of Graz, Universitaetsplatz 2, 8010 Graz, Austria.
6Department of Biology, University of Graz, Universitaetsplatz 2, 8010 Graz, Austria.
2026 - Volume: 66 Issue: 3 pages: 698-704
https://doi.org/10.24349/kki4-ccx2Historical taxonomic literature forms a fundamental basis of acarological research. Species descriptions, morphological diagnoses and faunistic surveys published during the 19th and 20th centuries remain essential for taxonomic research, nomenclatural stability and biodiversity studies under the International Code of Zoological Nomenclature (ICZN). Many original descriptions and revisions continue to represent the primary reference for systematic work. However, a substantial proportion of historical acarological publications have not been digitally indexed and remain accessible only in analogue collections. Access to primary taxonomic sources therefore often depends on physical library holdings, limiting discoverability and integration into contemporary biodiversity informatics workflows. Historically, bibliographic card index systems have served as structured tools for organizing scientific literature. Such collections frequently represent curated expert knowledge accumulated over decades and reflect specific thematic and taxonomic priorities. Their analogue format, however, restricts systematic querying and large-scale quantitative assessment. The Schuster Literature Collection represents one such curated archive with a strong focus on Acari. Compiled over several decades by Reinhart Schuster, a soil zoologist whose contributions significantly influenced Central European acarology (Krisper et al. 2023), the collection contains more than six thousand indexed publications spanning over two centuries. It includes taxonomic monographs, species descriptions, revisions and faunistic as well as ecological studies, with a particular emphasis on Oribatida. Until recently, the contents of this collection could only be accessed through manual consultation of the physical card index at the Department of Biology, University of Graz.
In line with broader concepts of taxonomic digitization, the goal is not merely to reproduce analogue structures in digital form, but to transform them into searchable and reusable information resources (Lyal 2016). Comparable efforts to render historical taxonomic knowledge digitally accessible exist at different scales. The Biodiversity Heritage Library (BHL– https://www.biodiversitylibrary.org/
) pursues large-scale digitization of legacy taxonomic literature internationally, combining OCR-based full-text access with automated taxonomic name recognition to link content across the corpus (Gwinn & Rinaldo 2009). At a regional scale directly comparable to the present context, ZOBODAT (Zoologisch-Botanische Datenbank – https://www.zobodat.at/
), hosted at the Biologiezentrum of the Oberösterreichisches Landesmuseum, has since 2005 digitized several million pages of Austrian and Central European zoological, botanical and geoscientific literature alongside biographical records of naturalists, demonstrating the long-term value of systematically indexing regional natural history literature (Gusenleitner & Malicky 2015). A related but methodologically distinct approach is followed by platforms such as DigiVol (DIGIVOL – https://volunteer.ala.org.au/
), which rely on crowdsourced transcription to convert analogue, label-based specimen metadata into structured digital records. The digitization of the Schuster card index follows the same underlying logic as these initiatives, transforming a curated but analogue knowledge resource into a structured digital dataset, while remaining deliberately limited in scope to a single, well-defined bibliographic archive. The present digitization effort does not involve the digitization of the publications themselves but focuses exclusively on the card index as a bibliographic metadata source. The primary objective is to establish a structured overview of what is contained in the collection. By converting the analogue index cards into a searchable digital dataset, the collection becomes transparent and systematically assessable for the first time. Researchers can determine which publications are documented in the archive without requiring physical access, and can evaluate its thematic, temporal and taxonomic composition. The digitization of the underlying publications represents a separate and substantially more extensive task. Making the full texts digitally available will require future coordinated efforts, depending on copyright status, technical feasibility and institutional resources. The current project therefore constitutes a foundational step: it creates the necessary bibliographic infrastructure upon which further digitization and long-term accessibility initiatives can progressively build.
Within the Schuster Literature Collection each publication is represented by a standardized card (14.7 × 10.4 cm) containing machine-typed bibliographic metadata and occasionally handwritten annotations (Fig. 1). Core elements include author(s), year, title, journal/source and a unique collection number linking the entry to the corresponding physical publication. All cards were scanned using a Canon image FORMULA DR-C225 document scanner with an automatic document feeder (ADF) in duplex mode at a resolution of 300 dpi.
As reverse sides contained handwritten notes in only ~5% of cases and were unsuitable for automated extraction, only front-side scans were processed; reverse images were archived.
Five metadata fields (author, year, title, journal/source, collection number) were defined as extraction targets. Automated metadata extraction was conducted using a custom-trained processor implemented in Google Cloud Document AI (Google LLC 2025). Almost hundred cards were manually annotated for supervised model training. A separate independent test set comprising 89 manually annotated index cards, which were not used during training, was used for model evaluation. Extraction performance was assessed using precision, recall and F1-score for each metadata field. Evaluation metrics are summarized in Table 1. Batch processing and data export were automated using custom Python scripts (Python Software Foundation 2024). Post-processing, including clustering-based harmonization of author and journal names, was performed in OpenRefine (OpenRefine 2023). Validation was based on the unique collection number; records lacking valid identifiers were excluded. The final dataset was exported in spreadsheet and JSON format for structured querying and quantitative analysis of historical acarological literature.
Download as
Metadata field
Precision (%)
Recall (%)
F1-score (%)
Author
98.9
96.6
97.7
Year
100.0
100.0
100.0
Collection number
98.2
98.2
98.2
Journal
85.0
77.3
81.0
Title
81.0
71.9
76.2
Handwritten notes
23.1
10.5
14.5
Overall
88.4
84.0
86.2
A total of 17,625 image files were generated from approximately 9,000 index cards by scanning both sides of each card. While the collection comprised around 9,000 cards, many represented notes, cross-references or incomplete entries without a unique collection number. Only front-side images were processed for automated metadata extraction. Evaluation of the custom-trained document processor yielded an overall precision of 88.4%, recall of 84.0% and an F1-score of 86.2% (Table 1). Author names (F1 = 97.7%), publication year (100%) and collection numbers (98.2%) showed the highest extraction performance, whereas journal (81.0%) and title (76.2%) fields were extracted with lower accuracy. Handwritten annotations achieved an F1-score of only 14.5% and were therefore excluded from subsequent analyses. Following clustering-based harmonization and validation based on the unique collection number, only records with a valid collection number were retained, resulting in a final dataset of 6,318 bibliographic records. The validated dataset spans 208 years, from Fauna Boica (Schrank 1803) to recent South African oribatid descriptions (Hugo-Coetzee 2011). The temporal coverage therefore captures nearly the entire historical development of modern acarology. Publication frequency increases markedly during the second half of the 20th century, reflecting intensified global acarological research (Fig. 2). This pattern corresponds with broader historical trends in acarology. Analyses of global publication activity indicate that the 1970s and 1980s represented a particularly productive period for mite taxonomy, with exceptionally high rates of newly described species worldwide (Zhang 2014). The Schuster Literature Collection reflects this trend, recording 1,487 publications from the 1970s, the decade with the highest output in the dataset (Fig. 2). In contrast, the Zoological Record comparison dataset reaches its maximum in the 1980s. This difference likely reflects differences in database coverage and in the underlying publication corpus. In particular, Zoological Record coverage is incomplete for the early nineteenth century, and low early counts should therefore be interpreted cautiously.
Title-based keyword frequency analysis demonstrates a strong taxonomic emphasis: Acari (2,075 occurrences), mites (874), species (663), Acarina (663) and Oribatida (613) are the most frequent terms. Overall, 73% of all records relate directly to Acari. Within these Acari-related records, approximately 45% concern Oribatida and around 25% concern Parasitiformes, based on title keyword frequencies.
The predominance of taxonomic and systematic contributions is reflected by numerous species descriptions, revisions and identification keys, including treatments of Apoplophora (Niedbała 2001), Castrichovella mesoafricana (Wiśniewski & Hirschmann 1990), Erythraeus mirjavehi (Gabryś 1989), Serratoppia guanicola (Subías & Arillo 1996), Antarctic Nanorchestes species (Strandtmann 1981, 1982) and Gymnodamaeus species (Weigmann & Mourek 2008). Major revisionary works include Brachychthoniidae (Moritz 1976), Zerconopsis and Pediculaster (Rack 1976), and revisions of the Berlese collection (Niedbała 1991). Classical identification keys are represented by Balogh (1961).
Ecological studies constitute a consistent secondary component, including intertidal taxa such as Ameronothrus lineatus (Søvik & Leinaas 2003), cave-dwelling mites (Schwiebea cavernicola – Pax 1940; Speothrombium monoculata – Robaux 1972), salt marsh assemblages (Schulte & Weigmann 1977) and forest and peatland communities (Moritz 1963). Trophic and decomposition-related investigations are represented by Behan-Pelletier & Hill (1983), while parasitic and phoretic associations include Macrocheles robustulus (Costa 1966) and Orthohalarachne letalis (Popp 1961). Biogeographically, the collection covers all major continents, including Antarctic Nanorchestes antarcticus (Block & Sømme 1981), South American Scutacaridae (Mahunka 1968), African Euphthiracaroidea (Niedbała 1993) and Neodiscopoma (Marais & Theron 1986), tropical faunas from Madagascar and Vietnam (Balogh 1960; Balogh & Mahunka 1967), and West African Oribatida (Wallwork 1961).
Morphological and developmental research is well represented, including the ''Rhagidia organ'' (Willmann 1934), spermatophore morphology (Schuster & Schuster 1969; Trávníček 1979), cuticle structure in Nanorchestidae (Rounsevell & Greenslade 1988) and post-embryonic development of Oppia nitens (Sengbusch & Sengbusch 1970). Parasitological and symbiotic associations include Riccardoella oudemansi (Thor 1932), Ophionyssus natricis (Yunker 1956), early symbiosis studies (Oudemans 1903) and mite assemblages in ant nests (Mahunka 1977). The thematic structure confirms the central role of alpha-taxonomy in acarology, particularly within Oribatida, and illustrates the historical integration of taxonomy, ecology and biogeography. The collection also preserves numerous historically important publications that continue to be relevant for contemporary acarological research. Examples include Schuster's (1960) revision of the genus Epilohmannia, Thor's (1925) contribution to the phylogeny and early development of the Acari, Thor's (1928) work on the family Bdellidae, and Haumann's (1997) doctoral thesis on the phylogeny of primitive oribatid mites. Such works illustrate the scientific breadth of the collection and highlight its value as a historical reference archive.
Beyond its disciplinary implications, the workflow contributes to open science and data sustainability. By converting analogue card indices into machine-readable datasets, bibliographic knowledge becomes searchable, interoperable and reusable. Structured metadata enhance transparency in literature retrieval and support reproducible taxonomic research. Although the dataset contains curated metadata rather than full-text content, its standardized design facilitates future linkage to digital repositories and biodiversity informatics infrastructures in line with FAIR (Findable, Accessible, Interoperable, and Reusable) principles. Additionally, the bibliographic dataset can enhance the correction of errors in citations and the detection of fabricated citations both often created by chatbots (Walters & Wilder 2023).
Finally, the digitization preserves the scholarly legacy of Reinhart Schuster, whose decades-long bibliographic curation formed the intellectual basis of the collection (Krisper et al. 2023). The transformation of this expert archive into a structured digital resource ensures continued accessibility of historically accumulated acarological knowledge and represents a continuation of scientific tradition within an open and sustainable research framework.
The bibliographic metadata of the Schuster Literature Collection are accessible through an online search interface at: https://michaelakerschbaumer.github.io/SchusterLiterature/
. The interface allows users to query the collection by author, year, title and journal/source. The underlying JSON file is maintained as a continuously updated working dataset, with corrections and metadata improvements incorporated over time. The GitHub Pages interface therefore provides access to the most recent version of the searchable catalogue. Stable versioned releases of the dataset are planned for future deposition in a long-term repository. Publications themselves remain part of the physical Schuster Literature Collection at the Department of Biology, University of Graz. Researchers interested in specific publications identified through the catalogue may contact the corresponding author to inquire about access.
This work was supported by the Austrian Academy of Sciences (ÖAW) within the project ''The Pillars of Soil''. We dedicate this work to the memory of Prof. Dr. Reinhart Schuster, whose bibliographic curation formed the foundation of this collection. We thank the Department of Biology, University of Graz, for access to the archive.

