Replication Data for: From nineteenth-century letters to entities: "a NER pipeline for French correspondence and its methodological lessons" - Article for Digital Humanities Benelux Journal (doi:10.34934/DVN/HNA7QO)

View:

Part 1: Document Description
Part 2: Study Description
Part 3: Data Files Description
Part 4: Variable Description
Part 5: Other Study-Related Materials
Entire Codebook

Document Description

Citation

Title:

Replication Data for: From nineteenth-century letters to entities: "a NER pipeline for French correspondence and its methodological lessons" - Article for Digital Humanities Benelux Journal

Identification Number:

doi:10.34934/DVN/HNA7QO

Distributor:

Social Sciences and Digital Humanities Archive – SODHA

Date of Distribution:

2026-08-25

Version:

1

Bibliographic Citation:

Zuzana Černáková; Fien Messens; Tess Dejaeghere; Julie M. Birkholz, 2026, "Replication Data for: From nineteenth-century letters to entities: "a NER pipeline for French correspondence and its methodological lessons" - Article for Digital Humanities Benelux Journal", https://doi.org/10.34934/DVN/HNA7QO, Social Sciences and Digital Humanities Archive – SODHA, V1, UNF:6:MemhfA5Fl4s1tvLWf4VR2A== [fileUNF]

Study Description

Citation

Title:

Replication Data for: From nineteenth-century letters to entities: "a NER pipeline for French correspondence and its methodological lessons" - Article for Digital Humanities Benelux Journal

Identification Number:

doi:10.34934/DVN/HNA7QO

Authoring Entity:

Zuzana Černáková (Department of History, Ghent University; Digital Research Lab, KBR)

Fien Messens (Department of History, Ghent University; KBR)

Tess Dejaeghere (Ghent Centre for Digital Humanities; Language Translation and Technology Team, Ghent University)

Julie M. Birkholz (Ghent Centre for Digital Humanities; Department of History, Ghent University; Digital Research Lab, KBR)

Distributor:

Social Sciences and Digital Humanities Archive – SODHA

Access Authority:

Messens, Fien

Depositor:

Messens, Fien

Date of Deposit:

2026-06-22

Holdings Information:

https://doi.org/10.34934/DVN/HNA7QO

Study Scope

Keywords:

Arts and Humanities, Named Entity Recognition (NER), Natural Language Processing (NLP), Historical NLP

Abstract:

This dataset accompanies the article From nineteenth-century letters to entities: A Named Entity Recognition pipeline for French correspondence and its methodological lessons. It contains the input data, preprocessing scripts, analysis notebooks, evaluation outputs, and experimental results used to develop and evaluate Named Entity Recognition (NER) workflows for nineteenth-century French correspondence from the François-Joseph Navez corpus (KBR – Royal Library of Belgium). The dataset documents the complete experimental workflow, from the construction of a manually annotated gold-standard corpus to the evaluation of off-the-shelf spaCy models, a custom-trained spaCy model, and transformer-based models (CamemBERT, CamemBERTav2, D'AlemBERT and Europeana BERT). The repository is organised into four directories. Input contains the Label Studio annotation exports, the preprocessed gold-standard corpus, and the training, development and test splits used throughout the experiments. Notebooks provides the Jupyter notebooks implementing the preprocessing, training and evaluation workflows. Output contains the evaluation results of the spaCy experiments, including model predictions, overall and per-entity performance metrics, mismatch analyses and summary tables. Results_transformers contains the outputs of the transformer-based experiments, including hyperparameter optimisation, cross-validation results, learning curves, statistical significance tests, prediction files and qualitative error analyses. Together, these files provide the complete computational workflow and all intermediate and final outputs required to reproduce the analyses presented in the accompanying publication.

Methodology and Processing

Sources Statement

Data Access

Notes:

These files contain training and evaluation data derived from nineteenth-century letters from the Navez Project. To protect access conditions associated with the source material, the files are restricted. Researchers may request access for non-commercial scholarly use. Requests will be evaluated on a case-by-case basis.

Other Study Description Materials

File Description--f7093

File: combined_eval_scores3.tab

  • Number of cases: 24

  • No. of variables per record: 13

  • Type of File: text/tab-separated-values

Notes:

UNF:6:IHdDRnEBZ4B81/KMQBqVKA==

File Description--f7103

File: cv_results.tab

  • Number of cases: 20

  • No. of variables per record: 6

  • Type of File: text/tab-separated-values

Notes:

UNF:6:72/ohGHOKcR5yXwDUXrubA==

File Description--f7080

File: errors_camembertav2.tab

  • Number of cases: 159

  • No. of variables per record: 10

  • Type of File: text/tab-separated-values

Notes:

UNF:6:dX1zhbyrPPk1DreziVEOHw==

File Description--f7055

File: errors_camembert.tab

  • Number of cases: 529

  • No. of variables per record: 10

  • Type of File: text/tab-separated-values

Notes:

UNF:6:LKJL6zjeIoC7+p0mefydwg==

File Description--f7106

File: errors_dalembert.tab

  • Number of cases: 185

  • No. of variables per record: 10

  • Type of File: text/tab-separated-values

Notes:

UNF:6:+TqzKZgt8g2h9mleixtmLw==

File Description--f7048

File: errors_europeana.tab

  • Number of cases: 161

  • No. of variables per record: 10

  • Type of File: text/tab-separated-values

Notes:

UNF:6:r0qXPIXovg15bXXLwViguQ==

File Description--f7077

File: headline_counter_table.tab

  • Number of cases: 16

  • No. of variables per record: 12

  • Type of File: text/tab-separated-values

Notes:

UNF:6:2raSn3MgN38lZjzDY60qZA==

File Description--f7056

File: headline_cv_summary.tab

  • Number of cases: 4

  • No. of variables per record: 3

  • Type of File: text/tab-separated-values

Notes:

UNF:6:LwsPP4NHhKrvsotyYYTR+w==

File Description--f7107

File: hp_search.tab

  • Number of cases: 72

  • No. of variables per record: 6

  • Type of File: text/tab-separated-values

Notes:

UNF:6:zu0B79FNEke0HpLSL1WIKw==

File Description--f7098

File: learning_curve.tab

  • Number of cases: 4

  • No. of variables per record: 4

  • Type of File: text/tab-separated-values

Notes:

UNF:6:fXKPhq8vFAp9d+v7Jlmpsg==

File Description--f7073

File: length_stratified.tab

  • Number of cases: 16

  • No. of variables per record: 5

  • Type of File: text/tab-separated-values

Notes:

UNF:6:/TrpFwgbWWGaD2ZI1dFDKQ==

File Description--f7064

File: length_summary.tab

  • Number of cases: 9

  • No. of variables per record: 5

  • Type of File: text/tab-separated-values

Notes:

UNF:6:XKylVdddSgIVxmbYCdrnFg==

File Description--f7058

File: navez_spacy_custom_trained_evalall_no_dk-2.tab

  • Number of cases: 4

  • No. of variables per record: 1

  • Type of File: text/tab-separated-values

Notes:

UNF:6:yq7KQEdVC0k1T86KGkSn0g==

File Description--f7075

File: nervaluate_counters_overall.tab

  • Number of cases: 16

  • No. of variables per record: 14

  • Type of File: text/tab-separated-values

Notes:

UNF:6:/3uciuaa/yLyO27c+p2Cyg==

File Description--f7082

File: per_tag_test.tab

  • Number of cases: 128

  • No. of variables per record: 14

  • Type of File: text/tab-separated-values

Notes:

UNF:6:YzGwFRM6nYl9P3yPhkiCEg==

File Description--f7089

File: significance_tests.tab

  • Number of cases: 6

  • No. of variables per record: 6

  • Type of File: text/tab-separated-values

Notes:

UNF:6:9YJ3w4uTtbYmfNKxJTh1UA==

File Description--f7097

File: test_predictions_camembertav2.tab

  • Number of cases: 17

  • No. of variables per record: 3

  • Type of File: text/tab-separated-values

Notes:

UNF:6:6PxO9ELAf/ENQQAY98ekCQ==

File Description--f7078

File: test_predictions_camembert.tab

  • Number of cases: 17

  • No. of variables per record: 3

  • Type of File: text/tab-separated-values

Notes:

UNF:6:wwYLN6EuL27q8HvnSckZ9A==

File Description--f7088

File: test_predictions_dalembert.tab

  • Number of cases: 17

  • No. of variables per record: 3

  • Type of File: text/tab-separated-values

Notes:

UNF:6:QOnGSP2nZCOfl3Bdo7Wayw==

File Description--f7072

File: test_predictions_europeana.tab

  • Number of cases: 17

  • No. of variables per record: 3

  • Type of File: text/tab-separated-values

Notes:

UNF:6:47AUFPWKuqAoGujhisXcrQ==

File Description--f7084

File: test_scores_per_seed.tab

  • Number of cases: 48

  • No. of variables per record: 13

  • Type of File: text/tab-separated-values

Notes:

UNF:6:BULzL9VsIcllUNnVwyPBzg==

Variable Description

List of Variables:

Variables

scenario

f7093 Location:

Variable Format: character

Notes: UNF:6:guT9fyOClBsdG7Zs4lSk0A==

workflow

f7093 Location:

Summary Statistics: StDev 0.7677718959499145; Max. 3.0; Mean 1.8; Valid 20.0; Min. 1.0

Variable Format: numeric

Notes: UNF:6:eeMzP9x4WHPanAuxyKNpAw==

model

f7093 Location:

Variable Format: character

Notes: UNF:6:d0yQGhpV/OO02tS5+gcRPQ==

correct

f7093 Location:

Summary Statistics: Max. 136.0; Valid 20.0; StDev 15.553896110316744; Min. 79.0; Mean 104.85

Variable Format: numeric

Notes: UNF:6:Blvpg/AXxXBYEc3IwdTPgg==

incorrect

f7093 Location:

Summary Statistics: Mean 33.5; Min. 0.0; StDev 25.736519210115254; Max. 69.0; Valid 20.0

Variable Format: numeric

Notes: UNF:6:UfJ7MHCo6S6kJPEg3QssYw==

partial

f7093 Location:

Summary Statistics: Valid 20.0; StDev 22.698191441981503; Mean 12.449999999999998; Max. 60.0; Min. 0.0;

Variable Format: numeric

Notes: UNF:6:Re+LPeoTtItmdi2Va3TWvg==

missed

f7093 Location:

Summary Statistics: Valid 20.0; Max. 67.0; StDev 6.75589411278462; Mean 57.2; Min. 49.0

Variable Format: numeric

Notes: UNF:6:U1MneR7DFJPh+EpR/WiUOw==

spurious

f7093 Location:

Summary Statistics: Max. 117.0; Valid 20.0; Mean 84.0; StDev 34.931511938589054; Min. 24.0

Variable Format: numeric

Notes: UNF:6:RANN6Xs+DsFn5pPPY5CQow==

possible

f7093 Location:

Summary Statistics: Mean 208.0; Valid 20.0; Min. 208.0; StDev 0.0; Max. 208.0

Variable Format: numeric

Notes: UNF:6:1jcKcIAFo9D5xsnGnMRXrQ==

actual

f7093 Location:

Summary Statistics: Min. 165.0; Mean 234.8; Max. 267.0; Valid 20.0; StDev 37.832873480472955;

Variable Format: numeric

Notes: UNF:6:0z0ZtH4pe55vvV7hqvvuvA==

p

f7093 Location:

Summary Statistics: Min. 0.3; Max. 0.8; Mean 0.4915; StDev 0.14180323654535124; Valid 20.0;

Variable Format: numeric

Notes: UNF:6:oVxpCCMQ7IZCSaEbqNJWGw==

r

f7093 Location:

Summary Statistics: Max. 0.65; StDev 0.08449696350819516; Valid 20.0; Min. 0.38; Mean 0.5335

Variable Format: numeric

Notes: UNF:6:tvCGrcWM4wbP/JjqhzPl+A==

f1

f7093 Location:

Summary Statistics: StDev 0.10457306687569926; Mean 0.5075; Min. 0.33; Valid 20.0; Max. 0.71

Variable Format: numeric

Notes: UNF:6:ev20S9tHslE2eGdfK2yWLQ==

model

f7103 Location:

Variable Format: character

Notes: UNF:6:58+a5/pKEXueFOgiC5eWsg==

fold

f7103 Location:

Summary Statistics: Mean 2.0; Max. 4.0; StDev 1.4509525002200232; Min. 0.0; Valid 20.0

Variable Format: numeric

Notes: UNF:6:GoMYq/48z1yzRJdm+KjV6A==

strict_f1

f7103 Location:

Summary Statistics: Max. 0.6378132118451024; Valid 20.0; StDev 0.10142737546736787; Mean 0.5034070979186959; Min. 0.3561128526645768

Variable Format: numeric

Notes: UNF:6:FlHLjD6/xq2jgsQy7D3CIA==

exact_f1

f7103 Location:

Summary Statistics: Mean 0.550975257485731; Valid 20.0; StDev 0.10674705669159901; Min. 0.3711598746081505; Max. 0.6788154897494305

Variable Format: numeric

Notes: UNF:6:0Lq2TriqZ3omdmbWgXOI2w==

partial_f1

f7103 Location:

Summary Statistics: Mean 0.6350428678276339; Max. 0.7653758542141231; StDev 0.11121765369336904; Min. 0.4369905956112852; Valid 20.0

Variable Format: numeric

Notes: UNF:6:Q1gc/LuFZJn9uq9ktw6uOw==

type_f1

f7103 Location:

Summary Statistics: Min. 0.46962616822429903; Mean 0.6204099419197274; Max. 0.7422360248447204; StDev 0.09864066985119124; Valid 20.0;

Variable Format: numeric

Notes: UNF:6:pw9c7yAtimEjW4enlILOJQ==

bucket

f7080 Location:

Variable Format: character

Notes: UNF:6:Z3MAwCntARuUewxy3T02sw==

ms

f7080 Location:

Variable Format: character

Notes: UNF:6:F1NvBj1Da9keQfhGu+N2Ww==

gold_type

f7080 Location:

Variable Format: character

Notes: UNF:6:zqFPcWFuMecwoZ/1J9OM1Q==

gold_start

f7080 Location:

Summary Statistics: Max. 875.0; Mean 245.99122807017534; StDev 243.9100550748577; Valid 114.0; Min. 0.0

Variable Format: numeric

Notes: UNF:6:Ia4QTUBFo2c5ta9NCKuTxA==

gold_end

f7080 Location:

Summary Statistics: Mean 247.32456140350888; Min. 0.0; Max. 875.0; StDev 243.7784608522042; Valid 114.0

Variable Format: numeric

Notes: UNF:6:zzh45T7/ua+2BmFtiRh8uA==

gold_text

f7080 Location:

Variable Format: character

Notes: UNF:6:9ndUGZB4m71YtiyapEhxzw==

pred_type

f7080 Location:

Variable Format: character

Notes: UNF:6:kf+pTrblULZLBOwxE8h+bQ==

pred_start

f7080 Location:

Summary Statistics: StDev 213.77687418212741; Mean 250.25555555555556; Valid 90.0; Max. 841.0; Min. 0.0;

Variable Format: numeric

Notes: UNF:6:o9WIrCkuY0IpbCNqmMI/gA==

pred_end

f7080 Location:

Summary Statistics: Valid 90.0; StDev 213.78318120070776; Max. 841.0; Mean 250.74444444444444; Min. 0.0;

Variable Format: numeric

Notes: UNF:6:lJiV+kwXTIuXOIHhYU+dPA==

pred_text

f7080 Location:

Variable Format: character

Notes: UNF:6:MP66zkrN14v/ht+snq6SRA==

bucket

f7055 Location:

Variable Format: character

Notes: UNF:6:KcDaCRK/7GRB1I9EUcO0jA==

ms

f7055 Location:

Variable Format: character

Notes: UNF:6:GSOKLOrg1oNrv9CiHAzHJQ==

gold_type

f7055 Location:

Variable Format: character

Notes: UNF:6:Bxn47JiR2JOwLLm/lziA3w==

gold_start

f7055 Location:

Summary Statistics: StDev 259.60811768527014; Max. 881.0; Min. 0.0; Mean 295.15151515151524; Valid 99.0

Variable Format: numeric

Notes: UNF:6:GoysvWcBtqd6mYNkqmWbtQ==

gold_end

f7055 Location:

Summary Statistics: Min. 1.0; StDev 259.4338090100685; Max. 885.0; Valid 99.0; Mean 296.43434343434325;

Variable Format: numeric

Notes: UNF:6:o9mFEoYa2oNVdbCrOtD+OA==

gold_text

f7055 Location:

Variable Format: character

Notes: UNF:6:xVV2UYufg/DmkW9lLJcylA==

pred_type

f7055 Location:

Variable Format: character

Notes: UNF:6:KuhNT2zgeJsB6Uwn46oqSg==

pred_start

f7055 Location:

Summary Statistics: Max. 931.0; StDev 209.12801729521792; Valid 522.0; Mean 259.5555555555561; Min. 0.0

Variable Format: numeric

Notes: UNF:6:UXjPSh7OWN90OQtbVQen3w==

pred_end

f7055 Location:

Summary Statistics: Max. 932.0; Valid 522.0; Min. 0.0; StDev 209.20831706344615; Mean 260.1609195402293

Variable Format: numeric

Notes: UNF:6:+qQhY+spzNcqnXJIuI36Rw==

pred_text

f7055 Location:

Variable Format: character

Notes: UNF:6:N5PyURGsNR1iDfbJPSr72w==

bucket

f7106 Location:

Variable Format: character

Notes: UNF:6:73tm7wEsvFfq3kSDYt7+Sw==

ms

f7106 Location:

Variable Format: character

Notes: UNF:6:gMwbm5A0T6rQPZoFNac92Q==

gold_type

f7106 Location:

Variable Format: character

Notes: UNF:6:yTuvHak0oGRJjLDuX4Ruww==

gold_start

f7106 Location:

Summary Statistics: Max. 875.0; Valid 118.0; Min. 0.0; Mean 253.4745762711865; StDev 233.1473878822464;

Variable Format: numeric

Notes: UNF:6:Jem/npq9hpaN0JiARU6toQ==

gold_end

f7106 Location:

Summary Statistics: Valid 118.0; Mean 254.83050847457636; StDev 233.01943373284368; Max. 875.0; Min. 1.0

Variable Format: numeric

Notes: UNF:6:+uzsdjTlBOknRVE2vUCYRA==

gold_text

f7106 Location:

Variable Format: character

Notes: UNF:6:0whxKQIn/X5ygDF2N5zx8g==

pred_type

f7106 Location:

Variable Format: character

Notes: UNF:6:mXbcxKB0DWCNNbaK3VgdLA==

pred_start

f7106 Location:

Summary Statistics: Max. 896.0; Mean 232.71774193548384; Valid 124.0; Min. 5.0; StDev 209.00645963814;

Variable Format: numeric

Notes: UNF:6:NpgVOTTT0x0stRR7+oCHQg==

pred_end

f7106 Location:

Summary Statistics: Mean 233.1290322580646; Valid 124.0; Min. 6.0; Max. 896.0; StDev 209.18210713878793;

Variable Format: numeric

Notes: UNF:6:sqsTt1vY1Dfqz/sbldvEaA==

pred_text

f7106 Location:

Variable Format: character

Notes: UNF:6:MVEsfjv9Ar7AfhD5/Dglxg==

bucket

f7048 Location:

Variable Format: character

Notes: UNF:6:PaeyOc/UW7LMF4KErQks6w==

ms

f7048 Location:

Variable Format: character

Notes: UNF:6:P7qH4V/ki4p8vpGmiE1haA==

gold_type

f7048 Location:

Variable Format: character

Notes: UNF:6:DqdaC91rdqw6ofiFDlQlzg==

gold_start

f7048 Location:

Summary Statistics: Min. 0.0; StDev 232.3513590405187; Max. 875.0; Mean 233.7672413793103; Valid 116.0;

Variable Format: numeric

Notes: UNF:6:AmW5b2qYS9IWLBD8CfPdvw==

gold_end

f7048 Location:

Summary Statistics: Valid 116.0; Min. 0.0; StDev 232.2088033125189; Max. 875.0; Mean 235.10344827586204;

Variable Format: numeric

Notes: UNF:6:vXqQ5az7uSFWmWZFftUzNQ==

gold_text

f7048 Location:

Variable Format: character

Notes: UNF:6:UBzxQx45p2SglkHccJqpGA==

pred_type

f7048 Location:

Variable Format: character

Notes: UNF:6:Noag+J/T3aIA5JUtr3OukA==

pred_start

f7048 Location:

Summary Statistics: Min. 2.0; Mean 231.1290322580646; Valid 93.0; Max. 852.0; StDev 198.8270619433125

Variable Format: numeric

Notes: UNF:6:1sPdzqobQbyWnaqkzCxLfQ==

pred_end

f7048 Location:

Summary Statistics: Valid 93.0; Min. 2.0; StDev 198.8528144010826; Mean 231.5806451612904; Max. 852.0

Variable Format: numeric

Notes: UNF:6:Z4gW3zWoU9kjCOjLVtMsJQ==

pred_text

f7048 Location:

Variable Format: character

Notes: UNF:6:WibaiZu9PYvHfdRi5KBxTg==

model

f7077 Location:

Variable Format: character

Notes: UNF:6:kpTfvbfk7/sGgBrZhUeuwQ==

scheme

f7077 Location:

Variable Format: character

Notes: UNF:6:hB9xBQ4Ec6uBrjXZ1F9prw==

correct

f7077 Location:

Summary Statistics: Max. 257.0; StDev 25.570816699250468; Mean 178.5; Min. 158.0; Valid 16.0

Variable Format: numeric

Notes: UNF:6:7LCZAzScqMmO/Ur9G8dMHg==

incorrect

f7077 Location:

Summary Statistics: Valid 16.0; Min. 0.0; Max. 96.0; StDev 33.378635881853135; Mean 38.5

Variable Format: numeric

Notes: UNF:6:mR6KSi/Lir/pyB659F3Grw==

partial

f7077 Location:

Summary Statistics: Mean 16.0; Valid 16.0; Max. 93.0; Min. 0.0; StDev 29.97332147093478;

Variable Format: numeric

Notes: UNF:6:wXR5OHCpo+p8ksbczGCi4g==

missed

f7077 Location:

Summary Statistics: Max. 51.0; Valid 16.0; Mean 36.0; Min. 2.0; StDev 20.668817092422103;

Variable Format: numeric

Notes: UNF:6:MnhZ0XKsgM12DO8gg1Y/aA==

spurious

f7077 Location:

Summary Statistics: Max. 458.0; StDev 171.78979403134906; Valid 16.0; Min. 61.0; Mean 170.5;

Variable Format: numeric

Notes: UNF:6:YAxnTAVvz+GP0jXi9fGA6A==

possible

f7077 Location:

Summary Statistics: Valid 16.0; StDev 0.0; Min. 269.0; Mean 269.0; Max. 269.0;

Variable Format: numeric

Notes: UNF:6:Zd3BbOJcUPPgBrdwgkQ/Jg==

actual

f7077 Location:

Summary Statistics: StDev 192.27549679214633; Min. 280.0; Mean 403.5; Max. 725.0; Valid 16.0;

Variable Format: numeric

Notes: UNF:6:Nq+MqFdbwx23DNebsr0U6Q==

precision

f7077 Location:

Summary Statistics: Mean 0.5279125; Valid 16.0; Min. 0.2359; StDev 0.15912954418753714; Max. 0.6966;

Variable Format: numeric

Notes: UNF:6:u924ES6xbvfa+zaVYS15tw==

recall

f7077 Location:

Summary Statistics: Mean 0.6933; Valid 16.0; Max. 0.9554; Min. 0.5874; StDev 0.10036617624146758

Variable Format: numeric

Notes: UNF:6:Zgyy3p/k9BLdi9xtUjfnrQ==

f1

f7077 Location:

Summary Statistics: Valid 16.0; StDev 0.12120868436983658; Min. 0.3441; Mean 0.5816125; Max. 0.7227;

Variable Format: numeric

Notes: UNF:6:qeDQxY2E1IwyvpSmqPQbhQ==

model

f7056 Location:

Variable Format: character

Notes: UNF:6:C1Eh2OlGL4uDPzqYdbUAEg==

cv_mean

f7056 Location:

Summary Statistics: Valid 4.0; Max. 0.5920197775226445; Min. 0.3710110533794434; StDev 0.10717392702937042; Mean 0.5034070979186958;

Variable Format: numeric

Notes: UNF:6:greybgZ9wXP+BT1ENd5cvQ==

cv_std

f7056 Location:

Summary Statistics: StDev 0.018774522899703643; Valid 4.0; Min. 0.00876821372661337; Max. 0.05197978066849702; Mean 0.03440492819439146

Variable Format: numeric

Notes: UNF:6:grHn+b35EE74Q0rm0w/IWg==

model

f7107 Location:

Variable Format: character

Notes: UNF:6:p5nC4LgNWL5cB7uA7OP3wQ==

lr

f7107 Location:

Summary Statistics: StDev 2.0331755499467835E-5; Mean 3.5833333333333335E-5; Valid 72.0; Min. 1.0E-5; Max. 8.0E-5

Variable Format: numeric

Notes: UNF:6:q9BhoZj9/N/iB3jlPk843g==

wd

f7107 Location:

Summary Statistics: Min. 0.0; Mean 0.005; Valid 72.0; StDev 0.005035088149780136; Max. 0.01;

Variable Format: numeric

Notes: UNF:6:FIb9pp4k2q89sVYUGmeWPA==

epochs

f7107 Location:

Summary Statistics: Mean 7.666666666666667; Min. 5.0; Valid 72.0; StDev 2.069224526445854; Max. 10.0;

Variable Format: numeric

Notes: UNF:6:y3+jqDE1GCAbyljL1nWHDw==

dev_strict_f1

f7107 Location:

Summary Statistics: Valid 72.0; StDev 0.16179550119195255; Mean 0.3494407811435388; Min. 0.0; Max. 0.5603217158176944

Variable Format: numeric

Notes: UNF:6:6hIZIx8dU4nF1n4K5ewi/g==

dev_partial_f1

f7107 Location:

Summary Statistics: Valid 72.0; StDev 0.19767326842942148; Max. 0.7131367292225201; Min. 0.0; Mean 0.4874900606173867

Variable Format: numeric

Notes: UNF:6:AXxtGlcep36N17NgCsw0hg==

fraction

f7098 Location:

Summary Statistics: Max. 1.0; Valid 4.0; Min. 0.25; Mean 0.625; StDev 0.3227486121839514

Variable Format: numeric

Notes: UNF:6:DLZRzcOjKc8xswdgXBmAHA==

n_letters

f7098 Location:

Summary Statistics: Min. 26.0; Max. 104.0; StDev 33.56585566713095; Mean 65.0; Valid 4.0

Variable Format: numeric

Notes: UNF:6:ZGH9W5nVkNFTmOJUgPrqoA==

dev_strict_f1

f7098 Location:

Summary Statistics: Max. 0.3578947368421053; Min. 0.2618351841028638; Mean 0.3086546513594165; StDev 0.041012080996138206; Valid 4.0

Variable Format: numeric

Notes: UNF:6:kJQ70Pcc4VJX7tgqMTfVCw==

test_strict_f1

f7098 Location:

Summary Statistics: Mean 0.29050894490237267; StDev 0.03536753183573118; Valid 4.0; Min. 0.25136612021857924; Max. 0.3369890329012961;

Variable Format: numeric

Notes: UNF:6:5RL/vXaVfOVSuXO8gRgSIg==

model

f7073 Location:

Variable Format: character

Notes: UNF:6:CQC7U1vPPyLCg9YVfWF6wA==

quartile

f7073 Location:

Variable Format: character

Notes: UNF:6:SWQcFbZTCpN/MxnBK8qqng==

n_letters

f7073 Location:

Summary Statistics: Valid 16.0; StDev 0.4472135954999579; Max. 5.0; Min. 4.0; Mean 4.25

Variable Format: numeric

Notes: UNF:6:GtAejoGVV+WDhTs4JdbXPQ==

strict_f1

f7073 Location:

Summary Statistics: StDev 0.12956119373002548; Mean 0.48339628265171414; Max. 0.6712328767123288; Valid 16.0; Min. 0.2238805970149254;

Variable Format: numeric

Notes: UNF:6:bSdHz4OCNBA8CYjwH2c7CA==

partial_f1

f7073 Location:

Summary Statistics: Valid 16.0; Mean 0.6104844560448139; Max. 0.7733333333333334; StDev 0.13181334952254545; Min. 0.34328358208955223;

Variable Format: numeric

Notes: UNF:6:j2rQ+f7glzbNHRtVz63mnA==

model

f7064 Location:

Variable Format: character

Notes: UNF:6:I/dhne68SAk/z1o+6U6Jdg==

split

f7064 Location:

Variable Format: character

Notes: UNF:6:8Dqbf2SbdKhIsYZb+93WCg==

mean_subwords

f7064 Location:

Summary Statistics: Mean 519.9942810457517; Max. 662.1176470588235; Min. 428.5882352941176; StDev 96.26966497755558; Valid 9.0;

Variable Format: numeric

Notes: UNF:6:XtLC2OQihvcSp/7PXyiuTA==

max_subwords

f7064 Location:

Summary Statistics: Max. 2583.0; Valid 9.0; StDev 640.2983245678879; Min. 999.0; Mean 1824.7777777777778;

Variable Format: numeric

Notes: UNF:6:9dGfZrvzOr2JpwXPZLY15w==

pct_over_512

f7064 Location:

Summary Statistics: StDev 5.883289508533153; Mean 37.14806435394671; Max. 47.05882352941176; Valid 9.0; Min. 29.807692307692307

Variable Format: numeric

Notes: UNF:6:m+ZZijwxapqFWZnrf9x3eA==

scenario;workflow;model;correct;incorrect;partial;missed;spurious;possible;actual;precision;recall;f1

f7058 Location:

Variable Format: character

Notes: UNF:6:yq7KQEdVC0k1T86KGkSn0g==

model

f7075 Location:

Variable Format: character

Notes: UNF:6:CQC7U1vPPyLCg9YVfWF6wA==

split

f7075 Location:

Variable Format: character

Notes: UNF:6:iRyMpghfxnLr3qVv5JG2ww==

tag

f7075 Location:

Variable Format: character

Notes: UNF:6:fkdvJ8NfYrQfObPiDmfmNg==

scheme

f7075 Location:

Variable Format: character

Notes: UNF:6:N0JYu8wqGOOIZ9OT7moybg==

correct

f7075 Location:

Summary Statistics: Min. 158.0; Max. 257.0; Mean 178.5; StDev 25.570816699250468; Valid 16.0

Variable Format: numeric

Notes: UNF:6:Dio7pLgF3ersydZYV19JNw==

incorrect

f7075 Location:

Summary Statistics: Mean 38.5; StDev 33.378635881853135; Valid 16.0; Min. 0.0; Max. 96.0

Variable Format: numeric

Notes: UNF:6:sRbO8gFJq9YPM7S8tRv5wg==

partial

f7075 Location:

Summary Statistics: Min. 0.0; Max. 93.0; StDev 29.97332147093478; Valid 16.0; Mean 16.0;

Variable Format: numeric

Notes: UNF:6:SXr8n+VFqWiZUGRNJ6NZfQ==

missed

f7075 Location:

Summary Statistics: StDev 20.668817092422103; Min. 2.0; Max. 51.0; Mean 36.0; Valid 16.0;

Variable Format: numeric

Notes: UNF:6:3kteUac9rkClfn30GDzahQ==

spurious

f7075 Location:

Summary Statistics: Max. 458.0; Mean 170.5; Valid 16.0; Min. 61.0; StDev 171.78979403134906;

Variable Format: numeric

Notes: UNF:6:H7feCSKfcQllDaXZIqYXAw==

possible

f7075 Location:

Summary Statistics: Min. 269.0; StDev 0.0; Valid 16.0; Max. 269.0; Mean 269.0

Variable Format: numeric

Notes: UNF:6:Zd3BbOJcUPPgBrdwgkQ/Jg==

actual

f7075 Location:

Summary Statistics: StDev 192.27549679214633; Max. 725.0; Mean 403.5; Valid 16.0; Min. 280.0;

Variable Format: numeric

Notes: UNF:6:N3gy9oU1HgdX4II7XHYCyg==

precision

f7075 Location:

Summary Statistics: Valid 16.0; Max. 0.696551724137931; Min. 0.23586206896551723; Mean 0.5279098326242723; StDev 0.15913042446406475

Variable Format: numeric

Notes: UNF:6:qMd3lRIFO1Zx4hSrXVlfzQ==

recall

f7075 Location:

Summary Statistics: StDev 0.10037174721189591; Valid 16.0; Min. 0.587360594795539; Mean 0.6933085501858736; Max. 0.9553903345724907

Variable Format: numeric

Notes: UNF:6:QXL3Aa+uDDS3+npf6h8tWA==

f1

f7075 Location:

Summary Statistics: Mean 0.5816155373025881; Valid 16.0; Min. 0.3440643863179075; Max. 0.7227191413237924; StDev 0.12121821060484483

Variable Format: numeric

Notes: UNF:6:qVMQmbS5B9sMzXqTIVKJWQ==

model

f7082 Location:

Variable Format: character

Notes: UNF:6:X5BN8XJ5YIjhihmTv4sR2g==

split

f7082 Location:

Variable Format: character

Notes: UNF:6:JLMKKUcUrzlFi0DbdlfGeQ==

tag

f7082 Location:

Variable Format: character

Notes: UNF:6:Q1VURVVLykE6y4/vQt2qgQ==

scheme

f7082 Location:

Variable Format: character

Notes: UNF:6:mAwlvWGuGWYXdPiXVjhAPQ==

correct

f7082 Location:

Summary Statistics: Min. 0.0; Max. 257.0; Valid 128.0; Mean 44.625; StDev 59.25522004315067

Variable Format: numeric

Notes: UNF:6:N3us5JBSOpzetds0ZTc0zA==

incorrect

f7082 Location:

Summary Statistics: Min. 0.0; Valid 128.0; Max. 96.0; StDev 17.61307674544527; Mean 9.625

Variable Format: numeric

Notes: UNF:6:sKNLVqo6o+U1tb05poJPlA==

partial

f7082 Location:

Summary Statistics: StDev 12.647243015990158; Mean 4.0; Max. 93.0; Min. 0.0; Valid 128.0;

Variable Format: numeric

Notes: UNF:6:IFKguEuxFMa8TR+bVhB2ug==

missed

f7082 Location:

Summary Statistics: Min. 0.0; Valid 128.0; Max. 51.0; StDev 13.1400691521329; Mean 9.0;

Variable Format: numeric

Notes: UNF:6:j0B3BWLXbcg62NKxCqMmYg==

spurious

f7082 Location:

Summary Statistics: Max. 458.0; Valid 128.0; StDev 82.08532146492453; Min. 0.0; Mean 42.625

Variable Format: numeric

Notes: UNF:6:pA8zRs5gYb4/m0E3C6Hplg==

possible

f7082 Location:

Summary Statistics: Max. 269.0; Min. 5.0; Valid 128.0; Mean 67.25; StDev 84.18464894569696;

Variable Format: numeric

Notes: UNF:6:D/OcIdlkuYp3buwj/F2rJQ==

actual

f7082 Location:

Summary Statistics: StDev 142.70585352367738; Valid 128.0; Mean 100.875; Min. 1.0; Max. 725.0;

Variable Format: numeric

Notes: UNF:6:0e7NAtaDqyNcldrExd/RuQ==

precision

f7082 Location:

Summary Statistics: StDev 0.25868873594147107; Min. 0.0; Mean 0.457882607416167; Valid 128.0; Max. 1.0

Variable Format: numeric

Notes: UNF:6:frXeISLH9ZBFHe1OMLmBuA==

recall

f7082 Location:

Summary Statistics: Min. 0.0; Max. 1.0; StDev 0.26905508190058236; Mean 0.5741175292008275; Valid 128.0

Variable Format: numeric

Notes: UNF:6:wvzElCF4yUEGe0E/ptr4dg==

f1

f7082 Location:

Summary Statistics: Mean 0.4720642287713962; Max. 0.8584070796460176; Min. 0.0; StDev 0.24287730198526605; Valid 128.0

Variable Format: numeric

Notes: UNF:6:j8qEaNwnfJFQ2nS6aFWRmA==

a

f7089 Location:

Variable Format: character

Notes: UNF:6:9adg+TH/qgHDspO5AOYyJw==

b

f7089 Location:

Variable Format: character

Notes: UNF:6:5BgcMwjSJpXvHMYYgUA2GA==

mean_diff

f7089 Location:

Summary Statistics: Max. -0.014253623730299734; Min. -0.23286558278850633; Mean -0.12106864964821236; Valid 6.0; StDev 0.10317161748169774

Variable Format: numeric

Notes: UNF:6:cfWL6ek4c4VcurE1jHkM1w==

ci95_lo

f7089 Location:

Summary Statistics: Mean -0.15791745353253395; Valid 6.0; Max. -0.037347367799293374; StDev 0.11358703502589378; Min. -0.27599502851993896;

Variable Format: numeric

Notes: UNF:6:wvON5infQ75SobPbpuO65A==

ci95_hi

f7089 Location:

Summary Statistics: Min. -0.17315830891823278; Max. 0.010631568937260551; Valid 6.0; StDev 0.085110869108066; Mean -0.07528920974788772;

Variable Format: numeric

Notes: UNF:6:P6ZCuLLlf//3GDyQ7Li3Pw==

p_2sided

f7089 Location:

Summary Statistics: Valid 6.0; StDev 0.1095694604653444; Mean 0.05766666666666667; Min. 0.0; Max. 0.274;

Variable Format: numeric

Notes: UNF:6:pHciBQhg68ueWWkX5mjJoA==

ms

f7097 Location:

Variable Format: character

Notes: UNF:6:MVYbDZTCwKmFIiA58Vxx5A==

true_iob

f7097 Location:

Variable Format: character

Notes: UNF:6:PeolXhxV6ezA6LWOB05rFg==

pred_iob

f7097 Location:

Variable Format: character

Notes: UNF:6:lCLbDgK/6vfjzu3jQdEHIQ==

ms

f7078 Location:

Variable Format: character

Notes: UNF:6:MVYbDZTCwKmFIiA58Vxx5A==

true_iob

f7078 Location:

Variable Format: character

Notes: UNF:6:PeolXhxV6ezA6LWOB05rFg==

pred_iob

f7078 Location:

Variable Format: character

Notes: UNF:6:BarU6CyBJtiTJm/pH1TFuQ==

ms

f7088 Location:

Variable Format: character

Notes: UNF:6:MVYbDZTCwKmFIiA58Vxx5A==

true_iob

f7088 Location:

Variable Format: character

Notes: UNF:6:PeolXhxV6ezA6LWOB05rFg==

pred_iob

f7088 Location:

Variable Format: character

Notes: UNF:6:iCvgRAWB7JPtKj1WqMwudA==

ms

f7072 Location:

Variable Format: character

Notes: UNF:6:MVYbDZTCwKmFIiA58Vxx5A==

true_iob

f7072 Location:

Variable Format: character

Notes: UNF:6:PeolXhxV6ezA6LWOB05rFg==

pred_iob

f7072 Location:

Variable Format: character

Notes: UNF:6:ha1yGA/ct/ODcspLhPicPg==

model

f7084 Location:

Variable Format: character

Notes: UNF:6:wBX2UHiWVwwDugXRAFGw/w==

seed

f7084 Location:

Summary Statistics: Valid 48.0; Max. 2024.0; Min. 42.0; StDev 925.5206546963191; Mean 729.6666666666667

Variable Format: numeric

Notes: UNF:6:1a3/v7xiN5JkO6ChFGSvzQ==

scheme

f7084 Location:

Variable Format: character

Notes: UNF:6:3hi6QXy5b0megqCl646e/A==

correct

f7084 Location:

Summary Statistics: Valid 48.0; StDev 24.7998280564847; Mean 177.22916666666666; Min. 151.0; Max. 257.0

Variable Format: numeric

Notes: UNF:6:GUoKMTE3c8DxtB93Uq8jyA==

incorrect

f7084 Location:

Summary Statistics: Max. 102.0; Mean 40.45833333333333; Valid 48.0; Min. 0.0; StDev 33.8802385967969

Variable Format: numeric

Notes: UNF:6:7U1iBKSkT8ecmvrQko8IJw==

partial

f7084 Location:

Summary Statistics: Min. 0.0; Valid 48.0; Max. 100.0; Mean 16.3125; StDev 30.062457148197883;

Variable Format: numeric

Notes: UNF:6:yFVwbAVnWkqYQBS83nHyTQ==

missed

f7084 Location:

Summary Statistics: StDev 19.254178877722083; Min. 2.0; Valid 48.0; Max. 51.0; Mean 35.0

Variable Format: numeric

Notes: UNF:6:FGOCkp0MFVokYc8exdpSNA==

spurious

f7084 Location:

Summary Statistics: Max. 503.0; Min. 61.0; Valid 48.0; Mean 173.25; StDev 178.23108356551103

Variable Format: numeric

Notes: UNF:6:vE89mIQe+NOqwbDfctq1GA==

possible

f7084 Location:

Summary Statistics: Valid 48.0; Mean 269.0; Min. 269.0; StDev 0.0; Max. 269.0

Variable Format: numeric

Notes: UNF:6:go4x4422IOs3AwH5YqYj0A==

actual

f7084 Location:

Summary Statistics: Valid 48.0; Min. 280.0; Mean 407.25; Max. 769.0; StDev 197.17154211821895

Variable Format: numeric

Notes: UNF:6:ja92tg6ekM5jfs+23CCyaQ==

precision

f7084 Location:

Summary Statistics: StDev 0.1614041161638638; Mean 0.525891676189172; Min. 0.21326397919375814; Max. 0.7295373665480427; Valid 48.0

Variable Format: numeric

Notes: UNF:6:VgvhfGomm/8BEBQnNfQvyw==

recall

f7084 Location:

Summary Statistics: Min. 0.5613382899628253; Mean 0.6891651177199505; Max. 0.9553903345724907; StDev 0.09838663123574758; Valid 48.0;

Variable Format: numeric

Notes: UNF:6:0EcO3YLGW1FXvrf23hFmzQ==

f1

f7084 Location:

Summary Statistics: StDev 0.1254759793685271; Max. 0.7454545454545454; Mean 0.578250574832892; Valid 48.0; Min. 0.31599229287090563

Variable Format: numeric

Notes: UNF:6:wqujaaVc+rWC559pknk/tQ==

Other Study-Related Materials

Label:

00_navez_gold_standard_preprocessing_json2df.ipynb

Text:

Preprocessing notebook that converts Label Studio JSON exports into the structured gold-standard corpus used throughout the experiments. The workflow merges annotation projects, extracts entity annotations, aligns character offsets, generates IOB labels, removes unsupported nested entities, and creates the training, development and test datasets used for the NER experiments.

Notes:

application/x-ipynb+json

Other Study-Related Materials

Label:

01a_navez_NER_spacy_off_the_shelf_no_dk_ner_pipe_sm.ipynb

Text:

Implements Workflow 1 described in the article by evaluating the off-the-shelf spaCy fr_core_news_sm model using only the isolated Named Entity Recognition (NER) component. The notebook evaluates the model on the held-out test set without additional domain-specific training.

Notes:

application/x-ipynb+json

Other Study-Related Materials

Label:

01b_navez_NER_spacy_off_the_shelf_no_dk_ner_pipe_lg.ipynb

Text:

Implements Workflow 1 using the larger fr_core_news_lg model. The notebook evaluates the isolated NER component on the historical correspondence corpus and compares its performance with the smaller spaCy model.

Notes:

application/x-ipynb+json

Other Study-Related Materials

Label:

02a_navez_NER_spacy_off_the_shelf_no_dk_nlp_pipe_sm.ipynb

Text:

Implements Workflow 2 by using the complete spaCy fr_core_news_sm pipeline, including all NLP components. The notebook assesses whether embedding the NER component within the full pipeline influences recognition performance.

Notes:

application/x-ipynb+json

Other Study-Related Materials

Label:

02b_navez_NER_spacy_off_the_shelf_no_dk_nlp_pipe_lg.ipynb

Text:

Implements Workflow 2 using the complete fr_core_news_lg pipeline. Results are compared with Workflow 1 to evaluate the effect of the full NLP pipeline on NER performance.

Notes:

application/x-ipynb+json

Other Study-Related Materials

Label:

03_navez_NER_spacy_custom-trained_dk.ipynb

Text:

Implements Workflow 3 described in the article. The notebook fine-tunes fr_core_news_lg on the Navez gold-standard corpus by extending the default spaCy model with the domain-specific entity types ART, EXH, GRP and LETT, in addition to the standard PER, LOC and ORG categories. It trains and evaluates the custom NER model on the historical correspondence corpus.

Notes:

application/x-ipynb+json

Other Study-Related Materials

Label:

best_hps.json

Text:

JSON file containing the final hyperparameter configuration selected for each transformer model (CamemBERT, CamemBERTav2, D'AlemBERT and Europeana BERT) after hyperparameter optimisation.

Notes:

application/json

Other Study-Related Materials

Label:

development_set-2.csv

Text:

Development (validation) partition of the corpus used for hyperparameter optimisation, model selection and intermediate evaluation during NER model training.

Notes:

text/csv

Other Study-Related Materials

Label:

development_set-2.pkl

Text:

Python Pickle version of the development dataset used in the training workflows.

Notes:

application/octet-stream

Other Study-Related Materials

Label:

headline_test_summary.csv

Text:

Summary of the final evaluation results obtained on the held-out test set for all transformer-based models averaged across three different seeds.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_gs_preprocessed-2.csv

Text:

Preprocessed version of the manually annotated corpus after cleaning and conversion from Label Studio. Contains token- and entity-level information used for model training and evaluation.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_gs_preprocessed-2.pkl

Text:

Python pickle version of the preprocessed gold-standard corpus, preserving data structures used directly in the notebooks.

Notes:

application/octet-stream

Other Study-Related Materials

Label:

Navez_ner_finetuning_weightedcrossentropy.ipynb

Text:

Notebook for fine-tuning and evaluating French transformer-based models for Named Entity Recognition on the Navez correspondence. It implements the experiments with CamemBERT, CamemBERTav2, D'AlemBERT and Europeana BERT, including weighted cross-entropy training, hyperparameter optimisation and evaluation using the Nervaluate framework.

Notes:

application/x-ipynb+json

Other Study-Related Materials

Label:

navez_spacy_custom_trained_evalall-2.csv

Text:

Overall evaluation metrics for the custom-trained spaCy model, including the four project-specific entity types introduced during fine-tuning. Reports precision, recall and F1 under the Nervaluate evaluation scenarios.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_custom_trained_evalpertype-2.csv

Text:

Detailed evaluation metrics for each entity type recognised by the custom-trained model, including both standard and domain-specific categories (PER, LOC, ORG, GRP, ART, EXH and LETT).

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_custom_trained_mismatches.csv

Text:

Token-level comparison between predicted and gold-standard IOB labels for the custom-trained spaCy model. Used for the qualitative error analysis discussed in the article.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_custom_trained_results.csv

Text:

Predictions generated by the custom-trained fr_core_news_lg model after fine-tuning on the Navez gold-standard corpus. Includes all project-specific entity types (PER, LOC, ORG, GRP, ART, EXH and LETT).

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_ner_lg_evalall-2.csv

Text:

Overall evaluation results for the off-the-shelf fr_core_news_lg model. Reports precision, recall and F1 scores under the Strict, Exact, Partial and Type evaluation scenarios using the Nervaluate framework.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_ner_lg_evalpertype-2.csv

Text:

Evaluation metrics for each entity type recognised by the off-the-shelf fr_core_news_lg model (PER, LOC and ORG). Enables comparison of model performance across entity categories.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_ner_lg_mismatches.csv

Text:

Token-level comparison between predicted and gold-standard IOB labels for the off-the-shelf fr_core_news_lg model. Used for qualitative error analysis and identification of recurring recognition errors.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_ner_lg_results.csv

Text:

Model predictions generated by the off-the-shelf fr_core_news_lg model using the isolated NER component (Workflow 1). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_ner_sm_evalall-2.csv

Text:

Overall evaluation results for the off-the-shelf fr_core_news_sm model. Reports precision, recall and F1 scores under the Strict, Exact, Partial and Type evaluation scenarios using the Nervaluate framework.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_ner_sm_evalpertype-2.csv

Text:

Evaluation metrics for each entity type recognised by the off-the-shelf fr_core_news_sm model (PER, LOC and ORG). Enables comparison of model performance across entity categories.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_ner_sm_mismatches.csv

Text:

Token-level comparison between predicted and gold-standard IOB labels for the off-the-shelf fr_core_news_sm model. Used for qualitative error analysis and identification of recurring recognition errors.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_ner_sm_results.csv

Text:

Model predictions generated by the off-the-shelf fr_core_news_sm model using the isolated NER component (Workflow 1). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_nlp_lg_evalall-2.csv

Text:

Overall evaluation results for the complete fr_core_news_lg NLP pipeline. Reports precision, recall and F1 scores under the Strict, Exact, Partial and Type evaluation scenarios using the Nervaluate framework.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_nlp_lg_evalpertype-2.csv

Text:

Evaluation metrics for each entity type recognised by the complete fr_core_news_lg pipeline (PER, LOC and ORG). Enables comparison with the isolated NER workflow and the custom-trained model.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_nlp_lg_mismatches.csv

Text:

Token-level comparison between predicted and gold-standard IOB labels for the complete fr_core_news_lg NLP pipeline. Used for qualitative error analysis and comparison with the isolated NER workflow.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_nlp_lg_results.csv

Text:

Model predictions generated by the off-the-shelf fr_core_news_lg model using the complete spaCy NLP pipeline (Workflow 2). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_nlp_sm_evalall-2.csv

Text:

Overall evaluation results for the complete fr_core_news_sm NLP pipeline. Reports precision, recall and F1 scores under the Strict, Exact, Partial and Type evaluation scenarios using the Nervaluate framework.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_nlp_sm_evalpertype-2.csv

Text:

Evaluation metrics for each entity type recognised by the complete fr_core_news_sm pipeline (PER, LOC and ORG). Enables comparison with the isolated NER workflow.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_nlp_sm_mismatches.csv

Text:

Token-level comparison between predicted and gold-standard IOB labels for the complete fr_core_news_sm NLP pipeline. Used for the qualitative error analysis and identification of recurring recognition errors.

Notes:

text/csv

Other Study-Related Materials

Label:

navez_spacy_off_the_shelf_nlp_sm_results.csv

Text:

Model predictions generated by the off-the-shelf fr_core_news_sm model using the complete spaCy NLP pipeline (Workflow 2). Contains predicted entities, spans, labels and corresponding gold-standard annotations for the held-out test set.

Notes:

text/csv

Other Study-Related Materials

Label:

project-237-at-2025-05-25-08-36-f874dc2a-2.json

Text:

Part 1 - JSON export from Label Studio containing manually annotated nineteenth-century French correspondence from the Navez Project. Includes entity annotations, transcriptions and project metadata. Used as source data for preprocessing and model training.

Notes:

application/json

Other Study-Related Materials

Label:

project-238-at-2025-05-25-08-34-694e26ec-2.json

Text:

Part 2 - JSON export from Label Studio containing manually annotated nineteenth-century French correspondence from the Navez Project. Includes entity annotations, transcriptions and project metadata. Used as source data for preprocessing and model training.

Notes:

application/json

Other Study-Related Materials

Label:

split_entity_counts.csv

Text:

Summary of the number of annotated entities per entity type in the training, development and test datasets. Used to document the composition of the gold-standard corpus and the experimental data splits.

Notes:

text/csv

Other Study-Related Materials

Label:

test_set-2.csv

Text:

Held-out test split used exclusively for final evaluation of the trained NER models.

Notes:

text/csv

Other Study-Related Materials

Label:

test_set-2.pkl

Text:

The file contains manually annotated letter transcriptions and their token-level entity labels. For each letter, it includes the manuscript identifier, full transcription, tokenized text, annotated entities and entity types, character and token spans, remapped annotations, and IOB tags.

Notes:

application/octet-stream

Other Study-Related Materials

Label:

training_set-2.csv

Text:

Held-out test partition of the gold-standard corpus used exclusively for final evaluation of the NER models after training.

Notes:

text/csv

Other Study-Related Materials

Label:

training_set-2.pkl

Text:

Python Pickle version of the held-out test dataset used in the evaluation pipeline.

Notes:

application/octet-stream