Integrating Collector and Author Roles Across Specimen and Publication Datasets.
Name
BISS_article_35866.pdf
Description
visibility:open
Size
62.39 KB
Format
Adobe PDF
Checksum (CRC64NVME)
mOP9sWNu0p4=
Resource type
Abstract
Date published
June 13, 2019
Abstract
This work builds on the outputs of a collector data-mining exercise applied to GBIF mobilised herbarium specimen metadata, which uses unsupervised learning (clustering) to identify collectors from minimal metadata associated with field collected specimens (the DarwinCore terms , and ). Here, we outline methods to integrate these data-mined collector entities (large scale dataset, aggregated from multiple sources, created programatically) with a dataset of author entities from the International Plant Names Index (smaller scale, single source dataset, created via editorial management). The integration process asserts a generic "scientist" entity with activities in different stages of the species description process: collecting and name publication. We present techniques to investigate specialisations including content - taxa of study - and activity stages: examining if individuals focus on collecting and/or name publication. Finally, we discuss generalisations of this initially herbarium-focussed data mining and record linkage process to enable applications in a wider context, particularly in zoological datasets.
Volume
3
Publisher
Pensoft Publishers
Place of publication
Sofia, Bulgaria
eISSN
2535-0897
Official URL
Related URL
Rights statement
In Copyright
Keywords
Additional information
This article is part of: ST08 - More than Names : Identifying and Crediting People in Biodiversity Data Edited by Simon Chagnoux, David Shorthouse, Anne Thessen.