Metrics
66,593 Downloads
Featured Dataverses

In order to use this feature you must have at least one published dataverse.

Publish Dataverse

Are you sure you want to publish your dataverse? Once you do so it must remain published.

Publish Dataverse

This dataverse cannot be published because the dataverse it is in has not been published.

Delete Dataverse

Are you sure you want to delete your dataverse? You cannot undelete this dataverse.

Advanced Search

11 to 20 of 243 Results
Comma Separated Values - 1.9 MB - MD5: 66e42633a3ae8e4f7ac2aa2bec9ba813
Token-level comparison between predicted and gold-standard IOB labels for the complete fr_core_news_lg NLP pipeline. Used for qualitative error analysis and comparison with the isolated NER workflow.
Comma Separated Values - 156.8 KB - MD5: 0c0b15b22c6286d52351ce6df21146e7
Model predictions generated by the off-the-shelf fr_core_news_lg model using the complete spaCy NLP pipeline (Workflow 2). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set.
Comma Separated Values - 2.1 MB - MD5: 25c9405590014091e5d08bf4aebb8301
Token-level comparison between predicted and gold-standard IOB labels for the complete fr_core_news_sm NLP pipeline. Used for the qualitative error analysis and identification of recurring recognition errors.
Comma Separated Values - 157.0 KB - MD5: 1128164075ad2b33e1902c44b764d000
Model predictions generated by the off-the-shelf fr_core_news_sm model using the complete spaCy NLP pipeline (Workflow 2). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set.
JSON - 198.6 KB - MD5: 72da0345995240cc733071556d1a647e
Part 1 - JSON export from Label Studio containing manually annotated nineteenth-century French correspondence from the Navez Project. Includes entity annotations, transcriptions and project metadata. Used as source data for preprocessing and model training.
JSON - 575.1 KB - MD5: 526f06d4f9ce02a351f068df4162d62b
Part 2 - JSON export from Label Studio containing manually annotated nineteenth-century French correspondence from the Navez Project. Includes entity annotations, transcriptions and project metadata. Used as source data for preprocessing and model training.
Comma Separated Values - 117.5 KB - MD5: c079a2cd638f0797e2b2ec4e06348e29
Held-out test split used exclusively for final evaluation of the trained NER models.
Unknown - 6.8 MB - MD5: ebeb69df35b8eecb39e9ce2342040c8d
The file contains manually annotated letter transcriptions and their token-level entity labels. For each letter, it includes the manuscript identifier, full transcription, tokenized text, annotated entities and entity types, character and token spans, remapped annotations, and IO...
Comma Separated Values - 730.5 KB - MD5: 8da5d01228c4818dddce6a97272b18be
Held-out test partition of the gold-standard corpus used exclusively for final evaluation of the NER models after training.
Unknown - 24.8 MB - MD5: 050192c781468337fb7718810e8316b7
Python Pickle version of the held-out test dataset used in the evaluation pipeline.
Add Data

Sign up or log in to create a dataverse or add a dataset.

Share Dataverse

Share this dataverse on your favorite social media networks.

Link Dataverse
Reset Modifications

Are you sure you want to reset the selected metadata fields? If you do this, any customizations (hidden, required, optional) you have done will no longer appear.