Repository logo
Home
Research Outputs
Collections
Statistics
Shared Repository Homepage
  1. Home
  2. Cultural Heritage Shared Repository Service
  3. Royal Botanic Gardens, Kew
  4. Article
  5. A benchmark dataset of herbarium specimen images with label data.

A benchmark dataset of herbarium specimen images with label data.

Thumbnail Image
Download
Name

BDJ_article_31817.pdf

Description
visibility:open
Size

448.29 KB

Format

Adobe PDF

Checksum (CRC64NVME)

8hTu8RO8usY=

Resource type
Journal article
Creator (person)
Dillen, Mathias
ORCIDORCID logo
Groom, Quentin
ORCIDORCID logo
Chagnoux, Simon
ORCIDORCID logo
Güntsch, Anton
ORCIDORCID logo
Hardisty, Alex
ORCIDORCID logo
Haston, Elspeth
ORCIDORCID logo
Livermore, Laurence
ORCIDORCID logo
Runnel, Veljo
ORCIDORCID logo
Schulman, Leif
ORCIDORCID logo
Willemse, Luc
ORCIDORCID logo
Wu, Zhengzhe
Phillips, Sarah
ORCIDORCID logo
Date published
February 8, 2019
Abstract
More and more herbaria are digitising their collections. Images of specimens are made available online to facilitate access to them and allow extraction of information from them. Transcription of the data written on specimens is critical for general discoverability and enables incorporation into large aggregated research datasets. Different methods, such as crowdsourcing and artificial intelligence, are being developed to optimise transcription, but herbarium specimens pose difficulties in data extraction for many reasons. To provide developers of transcription methods with a means of optimisation, we have compiled a benchmark dataset of 1,800 herbarium specimen images with corresponding transcribed data. These images originate from nine different collections and include specimens that reflect the multiple potential obstacles that transcription methods may encounter, such as differences in language, text format (printed or handwritten), specimen age and nomenclatural type status. We are making these specimens available with a Creative Commons Zero licence waiver and with permanent online storage of the data. By doing this, we are minimising the obstacles to the use of these images for transcription training. This benchmark dataset of images may also be used where a defined and documented set of herbarium specimens is needed, such as for the extraction of morphological traits, handwriting recognition and colour analysis of specimens.
Funder
Funder nameAwards
European Commission, European Union
H2020 programme (ICEDIG: RIA 777483)
Journal title
Biodiversity Data Journal
Volume
7
Article number
e31817
Publisher
Pensoft Publishers
Place of publication
Sofia, Bulgaria
ISSN
1314-2836
eISSN
1314-2828
Date accepted
February 4, 2019
Official URL
https://doi.org/10.3897/bdj.7.e31817
Related URL
https://bdj.pensoft.net/article/31817/
Rights statement
In Copyright
Licence
https://creativecommons.org/licenses/by/4.0/
DOI
10.3897/bdj.7.e31817
Keywords
Digitization
Labels
Herbarium digitization
Herbarium specimens
Specimen labels
Label data
Managed by the British Library and supported by the AHRC

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Cookie settings
  • End User Agreement
  • About
  • Contact
  • Help
Repository logo COAR Notify