11 to 20 of 243 Results
Comma Separated Values - 1.9 MB -
MD5: 66e42633a3ae8e4f7ac2aa2bec9ba813
Token-level comparison between predicted and gold-standard IOB labels for the complete fr_core_news_lg NLP pipeline. Used for qualitative error analysis and comparison with the isolated NER workflow. |
Comma Separated Values - 156.8 KB -
MD5: 0c0b15b22c6286d52351ce6df21146e7
Model predictions generated by the off-the-shelf fr_core_news_lg model using the complete spaCy NLP pipeline (Workflow 2). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set. |
Comma Separated Values - 2.1 MB -
MD5: 25c9405590014091e5d08bf4aebb8301
Token-level comparison between predicted and gold-standard IOB labels for the complete fr_core_news_sm NLP pipeline. Used for the qualitative error analysis and identification of recurring recognition errors. |
Comma Separated Values - 157.0 KB -
MD5: 1128164075ad2b33e1902c44b764d000
Model predictions generated by the off-the-shelf fr_core_news_sm model using the complete spaCy NLP pipeline (Workflow 2). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set. |
JSON - 198.6 KB -
MD5: 72da0345995240cc733071556d1a647e
Part 1 - JSON export from Label Studio containing manually annotated nineteenth-century French correspondence from the Navez Project. Includes entity annotations, transcriptions and project metadata. Used as source data for preprocessing and model training. |
JSON - 575.1 KB -
MD5: 526f06d4f9ce02a351f068df4162d62b
Part 2 - JSON export from Label Studio containing manually annotated nineteenth-century French correspondence from the Navez Project. Includes entity annotations, transcriptions and project metadata. Used as source data for preprocessing and model training. |
Comma Separated Values - 117.5 KB -
MD5: c079a2cd638f0797e2b2ec4e06348e29
Held-out test split used exclusively for final evaluation of the trained NER models. |
Unknown - 6.8 MB -
MD5: ebeb69df35b8eecb39e9ce2342040c8d
The file contains manually annotated letter transcriptions and their token-level entity labels. For each letter, it includes the manuscript identifier, full transcription, tokenized text, annotated entities and entity types, character and token spans, remapped annotations, and IO... |
Comma Separated Values - 730.5 KB -
MD5: 8da5d01228c4818dddce6a97272b18be
Held-out test partition of the gold-standard corpus used exclusively for final evaluation of the NER models after training. |
Unknown - 24.8 MB -
MD5: 050192c781468337fb7718810e8316b7
Python Pickle version of the held-out test dataset used in the evaluation pipeline. |
