Metrics
66,593 Downloads
Welcome to SODHA, the Belgian federal data archive for social sciences and the digital humanities!

Here you can find and deposit social science and digital humanities data in reusable form. Published datasets receive a DOI, making them citable like other types of publications. SODHA promotes open data by enabling reuse of research data and by safely preserving datasets in the long term.

SODHA is the Belgian service provider in the Consortium of European Social Science Data Archives (CESSDA) and is hosted by the State Archives of Belgium. SODHA was built with the help of DEMO (UCLouvain) and Interface Demography (VUB).

You can consult the SODHA Guide here, and you can read our policies here.

Want to learn more about SODHA? Consult our brochure or our presentation on the State Archives' website.

If you have any question, you can contact us at sodha@arch.be.
Featured Dataverses

In order to use this feature you must have at least one published dataverse.

Publish Dataverse

Are you sure you want to publish your dataverse? Once you do so it must remain published.

Publish Dataverse

This dataverse cannot be published because the dataverse it is in has not been published.

Delete Dataverse

Are you sure you want to delete your dataverse? You cannot undelete this dataverse.

Advanced Search

61 to 70 of 4,146 Results
Tabular Data - 12.5 KB - 14 Variables, 128 Observations - UNF:6:YzGwFRM6nYl9P3yPhkiCEg==
Evaluation results for each transformer-based model and each entity type (PER, LOC, ORG, GRP, ART, EXH and LETT) on the held-out test set under the Nervaluate scenarios. It enables detailed comparison of model performance across annotation categories.
JSON - 198.6 KB - MD5: 72da0345995240cc733071556d1a647e
Part 1 - JSON export from Label Studio containing manually annotated nineteenth-century French correspondence from the Navez Project. Includes entity annotations, transcriptions and project metadata. Used as source data for preprocessing and model training.
JSON - 575.1 KB - MD5: 526f06d4f9ce02a351f068df4162d62b
Part 2 - JSON export from Label Studio containing manually annotated nineteenth-century French correspondence from the Navez Project. Includes entity annotations, transcriptions and project metadata. Used as source data for preprocessing and model training.
Tabular Data - 559 B - 6 Variables, 6 Observations - UNF:6:9YJ3w4uTtbYmfNKxJTh1UA==
Results of statistical significance tests comparing the performance of the evaluated Named Entity Recognition models. Used to assess whether observed performance differences are statistically meaningful.
Comma Separated Values - 162 B - MD5: ae69733d2f22cb00b53e031a03f99e3f
Summary of the number of annotated entities per entity type in the training, development and test datasets. Used to document the composition of the gold-standard corpus and the experimental data splits.
Tabular Data - 96.0 KB - 3 Variables, 17 Observations - UNF:6:6PxO9ELAf/ENQQAY98ekCQ==
Named Entity Recognition predictions generated by the fine-tuned CamemBERTav2 model on the held-out test set. Includes predicted entity spans, labels and gold-standard annotations for model evaluation.
Tabular Data - 99.0 KB - 3 Variables, 17 Observations - UNF:6:wwYLN6EuL27q8HvnSckZ9A==
Named Entity Recognition predictions generated by the fine-tuned CamemBERT model on the held-out test set. Includes predicted entity spans, labels and gold-standard annotations for model evaluation.
Tabular Data - 96.3 KB - 3 Variables, 17 Observations - UNF:6:QOnGSP2nZCOfl3Bdo7Wayw==
Named Entity Recognition predictions generated by the fine-tuned D'AlemBERT model on the held-out test set. Includes predicted entity spans, labels and gold-standard annotations for model evaluation.
Tabular Data - 96.1 KB - 3 Variables, 17 Observations - UNF:6:47AUFPWKuqAoGujhisXcrQ==
Named Entity Recognition predictions generated by the fine-tuned Europeana BERT model on the held-out test set. Contains predicted entity spans, labels and corresponding gold-standard annotations used for the quantitative and qualitative evaluation of model performance.
Tabular Data - 4.9 KB - 13 Variables, 48 Observations - UNF:6:BULzL9VsIcllUNnVwyPBzg==
Evaluation results of the transformer-based Named Entity Recognition models across multiple random seeds. Reports the test-set performance of each training run and was used to assess the robustness and reproducibility of the experimental results by calculating mean performance ac...
Add Data

Sign up or log in to create a dataverse or add a dataset.

Share Dataverse

Share this dataverse on your favorite social media networks.

Link Dataverse
Reset Modifications

Are you sure you want to reset the selected metadata fields? If you do this, any customizations (hidden, required, optional) you have done will no longer appear.