1 to 10 of 243 Results
Comma Separated Values - 167.4 KB -
MD5: e2f0fbe2ba2975b0de463eb734e611ab
Development (validation) partition of the corpus used for hyperparameter optimisation, model selection and intermediate evaluation during NER model training. |
Unknown - 8.2 MB -
MD5: 8d104a218fd93d5b17d6bf1671104d10
Python Pickle version of the development dataset used in the training workflows. |
Comma Separated Values - 1015.3 KB -
MD5: 0469ba07b6aa2663770de059d36ef03f
Preprocessed version of the manually annotated corpus after cleaning and conversion from Label Studio. Contains token- and entity-level information used for model training and evaluation. |
Unknown - 33.0 MB -
MD5: 3a11affcc28d1dc3ac86f185c3f089a2
Python pickle version of the preprocessed gold-standard corpus, preserving data structures used directly in the notebooks. |
Comma Separated Values - 1.6 MB -
MD5: bd7431776ef54f320a6345ee2c39b5d0
Token-level comparison between predicted and gold-standard IOB labels for the custom-trained spaCy model. Used for the qualitative error analysis discussed in the article. |
Comma Separated Values - 155.9 KB -
MD5: 519958789ccf7a67f90669ce21103d90
Predictions generated by the custom-trained fr_core_news_lg model after fine-tuning on the Navez gold-standard corpus. Includes all project-specific entity types (PER, LOC, ORG, GRP, ART, EXH and LETT). |
Comma Separated Values - 2.0 MB -
MD5: 19e45bf3f0974a05524450f94b2b93da
Token-level comparison between predicted and gold-standard IOB labels for the off-the-shelf fr_core_news_lg model. Used for qualitative error analysis and identification of recurring recognition errors. |
Comma Separated Values - 156.9 KB -
MD5: 7c28c59d7b17430f781573f98a1a5833
Model predictions generated by the off-the-shelf fr_core_news_lg model using the isolated NER component (Workflow 1). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set. |
Comma Separated Values - 2.2 MB -
MD5: fdd354911d784ba25ace9972ad447335
Token-level comparison between predicted and gold-standard IOB labels for the off-the-shelf fr_core_news_sm model. Used for qualitative error analysis and identification of recurring recognition errors. |
Comma Separated Values - 157.2 KB -
MD5: 5f684b97b3c570eb8ada7751bdb6a075
Model predictions generated by the off-the-shelf fr_core_news_sm model using the isolated NER component (Workflow 1). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set. |
