This article was co-authored with generative AI. The proper nouns, figures, and history in the text have been cross-checked against public information where feasible, but the descriptions of historical background in particular may contain errors. Please check the primary sources at the end for accurate information. This article focuses mainly on the technical aspects of data construction.
This is the second in a series introducing landmark large-scale digital humanities projects. Related articles: Venice Time Machine / Stanford ORBIS / World Historical Gazetteer
What this project is
Sailing Letters is a project that rediscovered, digitized, and built into a corpus about 38,000 Dutch letters that had been seized from ships captured in war and long held in British public archives.
Many of them were private letters that never reached their addressees, recording the raw words of ordinary people — including women and children — who do not appear on the main stage of history. This article focuses mainly on the technical side of how this body of historical material was converted into structured data (a language corpus).
Background (why are there Dutch letters in Britain?)
Before getting into the technical story, let me cover just the minimum background.
In the 17th and 18th centuries, the Dutch Republic and England fought several wars (the Anglo-Dutch Wars). At the time, governments issued private ships a "letter of marque," legally authorizing the capture of enemy ships. Captured Dutch ships were brought to Britain's High Court of Admiralty, and for its review the papers aboard — ship's documents, cargo manifests, even private letters — were seized and held as evidence.
Thus letters that were originally meant to reach the Netherlands remained on the British side, and are now held at The National Archives in Kew, on the outskirts of London.
- This paragraph is a summary. The number of wars, the dates, and the details of the institutions are based on general historical accounts; for precise definitions, please consult specialist historical sources.
The story of rediscovery and digitization
- From the late 1970s to around 1980, the existence of this body of captured documents was (again) drawing attention from researchers on the Dutch side
- In 2005, on the initiative of the National Library of the Netherlands (KB, The Hague), the historian Roelof van Gelder created a preliminary inventory and estimated the letters from captured Dutch ships at about 38,000
- Of these, private letters number about 15,000, with the rest being commercial correspondence and the like
The KB launched this as the "Sailing Letters" undertaking, advancing imaging (through the preservation program Metamorfoze) and publishing a series of books for a general audience.
The technical core: building a language corpus
What put this project to use as scholarly data was Leiden University's research program Brieven als Buit / Letters as Loot. Led by the linguist Marijke van der Wal, it ran from 2008 to 2013.
From the standpoint of data construction, three points are important.
1. Diplomatic transcription
Rather than "correcting" the letters as they are input, the letters are transcribed faithfully as the writer actually wrote them — including misspellings, abbreviations, and punctuation. This is decisive for linguistic research, because the object of study is precisely "how ordinary people actually wrote." Normalizing to modern conventions would lose the most valuable information (the spelling, dialect, and errors of the time).
2. Detailed metadata design
For each letter, the writer's gender, social class, age, place of origin, relationship to the addressee, and so on are added as metadata. This structuring enables
- differences in language by social class, region, and gender
- distinguishing autograph from scribal writing (research into literacy)
and other quantitative analysis (sociolinguistics). This is the approach called "language history from below."
3. Citizen participation (crowdsourced transcription) and a public corpus
Part of the transcription work was carried out through volunteer citizen participation. The public corpus ultimately assembled (Brieven als Buit) comprises about 1,033 letters, searchable with metadata such as gender, class, and period.
Caution: "about 38,000 letters" is a rough estimate for the collection as a whole; what was precisely turned into a corpus and published is about 1,033 letters. Not everything has been digitized to the same precision.
Note: The site of the research institution hosting this corpus was temporarily inaccessible for a period due to the impact of a cyberattack in May 2026. If a linked page does not display, please try again after some time.
An easily confused separate project: the Prize Papers Project
Often confused with Sailing Letters is the German-led Prize Papers Project. Both handle the same body of documents at Kew, but they differ in scale, in who runs them, and in purpose.
| Sailing Letters / Brieven als Buit | Prize Papers Project | |
|---|---|---|
| Lead | The Netherlands (KB + Leiden University) | Directed by the Göttingen Academy of Sciences and Humanities / carried out by the University of Oldenburg (The National Archives holds and cooperates) |
| Scope | Only the Dutch letters among the captured documents | The entire Prize Papers collection (all languages, all nationalities) |
| Purpose | Precise corpus building for linguistics and social history | Comprehensive digitization and cataloging |
| Rough scale | Dutch letters, about 38,000 (public corpus about 1,033) | About 35,000 captured ships, over 160,000 undelivered letters, 19 languages, 3.5 million images (completion planned for 2037) |
The huge figures like "160,000 letters," "19 languages," and "3.5 million images" belong to the latter (the Prize Papers Project). They are a separate layer from the Sailing Letters Dutch corpus, so take care not to confuse them.
Significance from the standpoint of data construction
The methodological points this project demonstrated are as follows.
- The value of transcription without normalization: depending on the research aim, recording "faithfully as the original, errors and all" is essentially important — data design always follows "what do you want to analyze"
- Metadata determines the breadth of analysis: by structuring writer attributes in advance, a wide variety of quantitative analyses become possible later
- Combining citizen participation with expert research: a division of labor that gains scale through crowdsourced transcription while experts guarantee quality
Turning "letters that never arrived," a body of historical material, into a quantitatively analyzable corpus through diplomatic transcription and structured metadata — this design philosophy is instructive for any situation that turns historical text into data.




Comments
…