Metrics
66,593 Downloads
Welcome to SODHA, the Belgian federal data archive for social sciences and the digital humanities!

Here you can find and deposit social science and digital humanities data in reusable form. Published datasets receive a DOI, making them citable like other types of publications. SODHA promotes open data by enabling reuse of research data and by safely preserving datasets in the long term.

SODHA is the Belgian service provider in the Consortium of European Social Science Data Archives (CESSDA) and is hosted by the State Archives of Belgium. SODHA was built with the help of DEMO (UCLouvain) and Interface Demography (VUB).

You can consult the SODHA Guide here, and you can read our policies here.

Want to learn more about SODHA? Consult our brochure or our presentation on the State Archives' website.

If you have any question, you can contact us at sodha@arch.be.
Featured Dataverses

In order to use this feature you must have at least one published dataverse.

Publish Dataverse

Are you sure you want to publish your dataverse? Once you do so it must remain published.

Publish Dataverse

This dataverse cannot be published because the dataverse it is in has not been published.

Delete Dataverse

Are you sure you want to delete your dataverse? You cannot undelete this dataverse.

Advanced Search

41 to 50 of 4,343 Results
Jupyter Notebook - 4.6 MB - MD5: 8c2912db3ec1743bd4c24c1103a01c99
Notebook for fine-tuning and evaluating French transformer-based models for Named Entity Recognition on the Navez correspondence. It implements the experiments with CamemBERT, CamemBERTav2, D'AlemBERT and Europeana BERT, including weighted cross-entropy training, hyperparameter o...
Comma Separated Values - 424 B - MD5: 2431b25f16ced713f3f94efd19eede8d
Overall evaluation metrics for the custom-trained spaCy model, including the four project-specific entity types introduced during fine-tuning. Reports precision, recall and F1 under the Nervaluate evaluation scenarios.
Tabular Data - 462 B - 1 Variables, 4 Observations - UNF:6:yq7KQEdVC0k1T86KGkSn0g==
Overall evaluation of the custom-trained model restricted to the standard spaCy entity types (PER, LOC and ORG). This output enables direct comparison with the off-the-shelf spaCy models presented in the article.
Comma Separated Values - 2.0 KB - MD5: d770abff6f7b4ac2791c99f080b2ccc9
Detailed evaluation metrics for each entity type recognised by the custom-trained model, including both standard and domain-specific categories (PER, LOC, ORG, GRP, ART, EXH and LETT).
Comma Separated Values - 1.6 MB - MD5: bd7431776ef54f320a6345ee2c39b5d0
Token-level comparison between predicted and gold-standard IOB labels for the custom-trained spaCy model. Used for the qualitative error analysis discussed in the article.
Comma Separated Values - 155.9 KB - MD5: 519958789ccf7a67f90669ce21103d90
Predictions generated by the custom-trained fr_core_news_lg model after fine-tuning on the Navez gold-standard corpus. Includes all project-specific entity types (PER, LOC, ORG, GRP, ART, EXH and LETT).
Comma Separated Values - 410 B - MD5: 27ce89c83902be5851a3c2a3a4c0cf30
Overall evaluation results for the off-the-shelf fr_core_news_lg model. Reports precision, recall and F1 scores under the Strict, Exact, Partial and Type evaluation scenarios using the Nervaluate framework.
Comma Separated Values - 988 B - MD5: 95c489a1ea58b92b8d94ef795bc284ce
Evaluation metrics for each entity type recognised by the off-the-shelf fr_core_news_lg model (PER, LOC and ORG). Enables comparison of model performance across entity categories.
Comma Separated Values - 2.0 MB - MD5: 19e45bf3f0974a05524450f94b2b93da
Token-level comparison between predicted and gold-standard IOB labels for the off-the-shelf fr_core_news_lg model. Used for qualitative error analysis and identification of recurring recognition errors.
Comma Separated Values - 156.9 KB - MD5: 7c28c59d7b17430f781573f98a1a5833
Model predictions generated by the off-the-shelf fr_core_news_lg model using the isolated NER component (Workflow 1). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set.
Add Data

Sign up or log in to create a dataverse or add a dataset.

Share Dataverse

Share this dataverse on your favorite social media networks.

Link Dataverse
Reset Modifications

Are you sure you want to reset the selected metadata fields? If you do this, any customizations (hidden, required, optional) you have done will no longer appear.