Metrics
66,593 Downloads
Featured Dataverses

In order to use this feature you must have at least one published dataverse.

Publish Dataverse

Are you sure you want to publish your dataverse? Once you do so it must remain published.

Publish Dataverse

This dataverse cannot be published because the dataverse it is in has not been published.

Delete Dataverse

Are you sure you want to delete your dataverse? You cannot undelete this dataverse.

Advanced Search

1 to 10 of 243 Results
Comma Separated Values - 167.4 KB - MD5: e2f0fbe2ba2975b0de463eb734e611ab
Development (validation) partition of the corpus used for hyperparameter optimisation, model selection and intermediate evaluation during NER model training.
Unknown - 8.2 MB - MD5: 8d104a218fd93d5b17d6bf1671104d10
Python Pickle version of the development dataset used in the training workflows.
Comma Separated Values - 1015.3 KB - MD5: 0469ba07b6aa2663770de059d36ef03f
Preprocessed version of the manually annotated corpus after cleaning and conversion from Label Studio. Contains token- and entity-level information used for model training and evaluation.
Unknown - 33.0 MB - MD5: 3a11affcc28d1dc3ac86f185c3f089a2
Python pickle version of the preprocessed gold-standard corpus, preserving data structures used directly in the notebooks.
Comma Separated Values - 1.6 MB - MD5: bd7431776ef54f320a6345ee2c39b5d0
Token-level comparison between predicted and gold-standard IOB labels for the custom-trained spaCy model. Used for the qualitative error analysis discussed in the article.
Comma Separated Values - 155.9 KB - MD5: 519958789ccf7a67f90669ce21103d90
Predictions generated by the custom-trained fr_core_news_lg model after fine-tuning on the Navez gold-standard corpus. Includes all project-specific entity types (PER, LOC, ORG, GRP, ART, EXH and LETT).
Comma Separated Values - 2.0 MB - MD5: 19e45bf3f0974a05524450f94b2b93da
Token-level comparison between predicted and gold-standard IOB labels for the off-the-shelf fr_core_news_lg model. Used for qualitative error analysis and identification of recurring recognition errors.
Comma Separated Values - 156.9 KB - MD5: 7c28c59d7b17430f781573f98a1a5833
Model predictions generated by the off-the-shelf fr_core_news_lg model using the isolated NER component (Workflow 1). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set.
Comma Separated Values - 2.2 MB - MD5: fdd354911d784ba25ace9972ad447335
Token-level comparison between predicted and gold-standard IOB labels for the off-the-shelf fr_core_news_sm model. Used for qualitative error analysis and identification of recurring recognition errors.
Comma Separated Values - 157.2 KB - MD5: 5f684b97b3c570eb8ada7751bdb6a075
Model predictions generated by the off-the-shelf fr_core_news_sm model using the isolated NER component (Workflow 1). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set.
Add Data

Sign up or log in to create a dataverse or add a dataset.

Share Dataverse

Share this dataverse on your favorite social media networks.

Link Dataverse
Reset Modifications

Are you sure you want to reset the selected metadata fields? If you do this, any customizations (hidden, required, optional) you have done will no longer appear.