Metrics
66,593 Downloads
Welcome to SODHA, the Belgian federal data archive for social sciences and the digital humanities!

Here you can find and deposit social science and digital humanities data in reusable form. Published datasets receive a DOI, making them citable like other types of publications. SODHA promotes open data by enabling reuse of research data and by safely preserving datasets in the long term.

SODHA is the Belgian service provider in the Consortium of European Social Science Data Archives (CESSDA) and is hosted by the State Archives of Belgium. SODHA was built with the help of DEMO (UCLouvain) and Interface Demography (VUB).

You can consult the SODHA Guide here, and you can read our policies here.

Want to learn more about SODHA? Consult our brochure or our presentation on the State Archives' website.

If you have any question, you can contact us at sodha@arch.be.
Featured Dataverses

In order to use this feature you must have at least one published dataverse.

Publish Dataverse

Are you sure you want to publish your dataverse? Once you do so it must remain published.

Publish Dataverse

This dataverse cannot be published because the dataverse it is in has not been published.

Delete Dataverse

Are you sure you want to delete your dataverse? You cannot undelete this dataverse.

Advanced Search

51 to 60 of 4,139 Results
Comma Separated Values - 157.2 KB - MD5: 5f684b97b3c570eb8ada7751bdb6a075
Model predictions generated by the off-the-shelf fr_core_news_sm model using the isolated NER component (Workflow 1). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set.
Comma Separated Values - 434 B - MD5: 0b166ca4796478707cf53976e9ed3da1
Overall evaluation results for the complete fr_core_news_lg NLP pipeline. Reports precision, recall and F1 scores under the Strict, Exact, Partial and Type evaluation scenarios using the Nervaluate framework.
Comma Separated Values - 1.0 KB - MD5: 02002572ce6487ca896a367b54f4d998
Evaluation metrics for each entity type recognised by the complete fr_core_news_lg pipeline (PER, LOC and ORG). Enables comparison with the isolated NER workflow and the custom-trained model.
Comma Separated Values - 1.9 MB - MD5: 66e42633a3ae8e4f7ac2aa2bec9ba813
Token-level comparison between predicted and gold-standard IOB labels for the complete fr_core_news_lg NLP pipeline. Used for qualitative error analysis and comparison with the isolated NER workflow.
Comma Separated Values - 156.8 KB - MD5: 0c0b15b22c6286d52351ce6df21146e7
Model predictions generated by the off-the-shelf fr_core_news_lg model using the complete spaCy NLP pipeline (Workflow 2). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set.
Comma Separated Values - 438 B - MD5: 60738aaa0921c5c4de7de9e63ecca08a
Overall evaluation results for the complete fr_core_news_sm NLP pipeline. Reports precision, recall and F1 scores under the Strict, Exact, Partial and Type evaluation scenarios using the Nervaluate framework.
Comma Separated Values - 1.0 KB - MD5: 0063ddb26e8f5f36183bac8e2cffb6d2
Evaluation metrics for each entity type recognised by the complete fr_core_news_sm pipeline (PER, LOC and ORG). Enables comparison with the isolated NER workflow.
Comma Separated Values - 2.1 MB - MD5: 25c9405590014091e5d08bf4aebb8301
Token-level comparison between predicted and gold-standard IOB labels for the complete fr_core_news_sm NLP pipeline. Used for the qualitative error analysis and identification of recurring recognition errors.
Comma Separated Values - 157.0 KB - MD5: 1128164075ad2b33e1902c44b764d000
Model predictions generated by the off-the-shelf fr_core_news_sm model using the complete spaCy NLP pipeline (Workflow 2). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set.
Tabular Data - 1.8 KB - 14 Variables, 16 Observations - UNF:6:/3uciuaa/yLyO27c+p2Cyg==
Aggregated Nervaluate statistics reporting the numbers of correct, incorrect, partial, missed and spurious entity predictions on the held-out test set under the different evaluation scenarios (Strict, Exact, Partial and Type) for all transformer-based models.
Add Data

Sign up or log in to create a dataverse or add a dataset.

Share Dataverse

Share this dataverse on your favorite social media networks.

Link Dataverse
Reset Modifications

Are you sure you want to reset the selected metadata fields? If you do this, any customizations (hidden, required, optional) you have done will no longer appear.