11 to 20 of 4,336 Results
Adobe PDF - 264.7 KB -
MD5: 934208dab49bf486ca3b81f6c34f44c6
Stata export of the dataset v 2.0 |
Tabular Data - 8.3 MB - 224 Variables, 2254 Observations - UNF:6:2wptholm/ZhVLRkyMklF/w==
CSV export of the dataset v 2.0 |
Tabular Data - 1.6 MB - 224 Variables, 2254 Observations - UNF:6:qb1kur9VxUc6vdWn1vPtfw==
SPSS export of the dataset v 2.0 |
Tabular Data - 1.9 MB - 224 Variables, 2254 Observations - UNF:6:qb1kur9VxUc6vdWn1vPtfw==
Stata export of the dataset v 2.0 |
Adobe PDF - 355.2 KB -
MD5: 86278cc92a3598d82e6eb8b5586fc2d7
Data user guide v 2.0 |
Aug 25, 2026
Zuzana Černáková; Fien Messens; Tess Dejaeghere; Julie M. Birkholz, 2026, "Replication Data for: From nineteenth-century letters to entities: "a NER pipeline for French correspondence and its methodological lessons" - Article for Digital Humanities Benelux Journal", https://doi.org/10.34934/DVN/HNA7QO, Social Sciences and Digital Humanities Archive – SODHA, V1, UNF:6:MemhfA5Fl4s1tvLWf4VR2A== [fileUNF]
This dataset accompanies the article From nineteenth-century letters to entities: A Named Entity Recognition pipeline for French correspondence and its methodological lessons. It contains the input data, preprocessing scripts, analysis notebooks, evaluation outputs, and experimen... |
Jupyter Notebook - 465.1 KB -
MD5: 6062b90f470e79ad4062aa678de087d3
Preprocessing notebook that converts Label Studio JSON exports into the structured gold-standard corpus used throughout the experiments. The workflow merges annotation projects, extracts entity annotations, aligns character offsets, generates IOB labels, removes unsupported neste... |
Jupyter Notebook - 201.0 KB -
MD5: fafce906de349b467636197fa64a4d95
Implements Workflow 1 described in the article by evaluating the off-the-shelf spaCy fr_core_news_sm model using only the isolated Named Entity Recognition (NER) component. The notebook evaluates the model on the held-out test set without additional domain-specific training. |
Jupyter Notebook - 248.3 KB -
MD5: 4d33ecdf3831757bac0319676a57bb2a
Implements Workflow 1 using the larger fr_core_news_lg model. The notebook evaluates the isolated NER component on the historical correspondence corpus and compares its performance with the smaller spaCy model. |
Jupyter Notebook - 200.0 KB -
MD5: 2998c19d3448da4d9ed7a7776de6e9dc
Implements Workflow 2 by using the complete spaCy fr_core_news_sm pipeline, including all NLP components. The notebook assesses whether embedding the NER component within the full pipeline influences recognition performance. |
