Digital Humanities Vocabulary in English

20 essential digital humanities words covering text mining, digitisation, archives and computational methods — ideal for B2–C1 learners studying humanities, library science, or information studies.

Pedagogically reviewed by LexFizz Team

Digital humanities (DH) sits at the intersection of technology and humanistic inquiry, applying computational tools to research in history, literature, linguistics, art history, and cultural studies. As universities invest in digital scholarship centres and libraries digitise millions of historical documents, the vocabulary of digital humanities has become increasingly important for researchers, librarians, archivists, and anyone engaged in the study or preservation of cultural heritage. Terms such as text mining, metadata, corpus, digitisation, and open access now appear in grant applications, research proposals, conference programmes, and job descriptions across the cultural sector, and their everyday senses can be checked against the Oxford Learner's Dictionaries. At B2–C1 level, mastering this vocabulary allows you to participate in interdisciplinary research teams, read DH scholarship, apply for positions in digital archives and libraries, and contribute to projects that are transforming how we understand and access the past. The field also raises important critical questions about bias in algorithms, unequal access to digital resources, and the long-term preservation of digital heritage — questions that require precise language to discuss effectively. This vocabulary bridges the worlds of traditional humanities scholarship and computing, making it uniquely valuable for learners who want to work at the cutting edge of how knowledge is created, shared, and preserved in the digital age.

Essential Digital Humanities Words

WordPronunciationPart of speechDefinitionExample sentence
digitisation/ˌdɪdʒ.ɪ.taɪˈzeɪ.ʃən/nounthe conversion of physical objects or documents into digital formats that can be stored, searched, and shared electronicallyThe library's digitisation project made 500,000 manuscript pages freely available online.
text mining/tekst ˈmaɪ.nɪŋ/noun phrasethe automated extraction of patterns, trends, and insights from large collections of text using computational methodsText mining of Victorian newspapers revealed a significant increase in the frequency of the word “progress” after 1850.
corpus/ˈkɔː.pəs/nouna structured, large-scale collection of texts assembled for research and computational analysisThe researchers compiled a corpus of 10,000 scientific papers to study how scientific language has changed over 200 years.
metadata/ˈmet.ə.deɪ.tə/nounstructured data that describes and provides information about other data, making digital objects discoverable and usableRich metadata including date, author, and provenance makes archived documents far easier for researchers to find and cite.
open access/ˈəʊ.pən ˈæk.ses/noun phrasethe free, unrestricted online availability of research outputs for anyone to read, download, and reuseThe journal switched to open access publication, making all articles freely available to researchers worldwide.
topic modelling/ˈtɒp.ɪk ˈmɒd.əl.ɪŋ/noun phrasea machine learning technique for identifying recurring themes or topics within a large collection of documentsTopic modelling of 18th-century pamphlets revealed the gradual emergence of concepts of individual rights and liberty.
distant reading/ˈdɪs.tənt ˈriː.dɪŋ/noun phraseFranco Moretti's term for the computational analysis of large literary corpora to identify patterns across many texts simultaneouslyDistant reading revealed that the average length of English novels increased significantly between 1750 and 1850.
digital archive/ˈdɪdʒ.ɪ.tl ˈɑː.kaɪv/noun phrasean online repository of digitised or born-digital materials preserved and made accessible for research and public useThe digital archive of wartime letters received over 100,000 visits in its first year online.
markup/ˈmɑːk.ʌp/nounthe encoding of text with tags that identify its structure, content, or features for computational processingTEI markup was applied to the manuscript to tag all named persons, places, and dates for searchability.
visualisation/ˌvɪʒ.u.ə.laɪˈzeɪ.ʃən/nounthe graphical representation of data or text to reveal patterns, relationships, and trends that are difficult to see in raw formThe visualisation of network connections between 18th-century correspondents revealed previously unknown intellectual communities.
crowdsourcing/ˈkraʊd.sɔː.sɪŋ/nounthe use of contributions from a large online community to complete tasks such as transcribing, tagging, or classifying materialsCrowdsourcing allowed thousands of volunteers to help transcribe handwritten census records in just six months.
optical character recognition/ˈɒp.tɪ.kəl ˈkær.ɪk.tər ˌrek.əɡˈnɪʃ.ən/noun phrasesoftware that converts scanned images of text into machine-readable characters that can be searched and editedOptical character recognition allowed researchers to full-text search a newspaper archive of two million pages.
provenance/ˈprɒv.ə.nəns/nounthe documented history of the ownership and origin of a manuscript, artefact, or documentThe provenance of the manuscript was traced back to a medieval monastery through a chain of documented sales.
interoperability/ˌɪn.tər.ɒp.ər.əˈbɪl.ɪ.ti/nounthe ability of different digital systems and collections to work together by sharing standards and formatsAdopting common metadata standards ensures interoperability between the archive and other international collections.
digital preservation/ˈdɪdʒ.ɪ.tl ˌprez.əˈveɪ.ʃən/noun phrasethe long-term management and maintenance of digital content to ensure its continued accessibility and integrity over timeDigital preservation requires ongoing migration to new formats as older file types become obsolete.
natural language processing/ˈnætʃ.ər.əl ˈlæŋ.ɡwɪdʒ ˈprəʊ.ses.ɪŋ/noun phrasea branch of artificial intelligence concerned with the computational analysis and generation of human languageNatural language processing tools were used to identify all instances of named entities across the digitised newspaper corpus.
data curation/ˈdeɪ.tə kjʊəˈreɪ.ʃən/noun phrasethe active management of research data throughout its lifecycle to ensure its quality, accessibility, and long-term usabilityData curation standards require researchers to document their datasets thoroughly before depositing them in a repository.
linked open data/lɪŋkt ˈəʊ.pən ˈdeɪ.tə/noun phrasea set of practices for publishing and connecting structured data on the web using open standards, enabling discovery across datasetsLinked open data standards allowed the museum's collection to be connected to relevant records in national and international databases.
algorithm/ˈæl.ɡə.rɪ.ðəm/nouna set of rules or instructions followed by a computer to solve a problem or analyse dataThe algorithm classified 50,000 documents by genre with over 85% accuracy compared to expert human judgment.
digital edition/ˈdɪdʒ.ɪ.tl ɪˈdɪʃ.ən/noun phrasea scholarly electronic publication of a historical text with features such as encoded markup, annotations, and image facsimilesThe digital edition of the medieval chronicle allows users to compare three manuscript versions side by side.

Practise with exercises

Related vocabulary topics

Explore all vocabulary topics

Discover hundreds of topic word lists with free interactive exercises on LexFizz.

All Vocabulary Topics

Frequently Asked Questions

What is digital humanities?

Digital humanities (DH) is an interdisciplinary field that applies computational tools and methods to humanities research in areas such as history, literature, linguistics, art history, and cultural studies. It encompasses digitising manuscripts, building digital archives, applying text mining to large corpora, creating visualisations of cultural data, and developing new forms of digital scholarly publishing. DH emerged in the 1940s with Father Busa's index of Thomas Aquinas's works and has grown substantially with the rise of the internet, large digitised collections, and open-source computational tools.

What is text mining in digital humanities?

Text mining (also called text analytics) is the automated process of extracting meaningful patterns, trends, and insights from large bodies of text using computational methods. In digital humanities, it is used to analyse corpora of thousands or millions of documents impossible to read manually. Techniques include topic modelling, sentiment analysis, word frequency analysis, and named entity recognition. Text mining enables ‘distant reading’ — the analysis of large-scale patterns across many texts simultaneously, complementing traditional close reading of individual texts.

What is digitisation?

Digitisation is the process of converting physical objects, documents, or media into digital formats that can be stored, searched, and shared electronically. In cultural heritage institutions, digitisation projects aim to improve access to collections, preserve fragile materials, and enable new research possibilities. Mass digitisation projects such as Google Books and the Internet Archive have made millions of texts freely searchable. Digitisation raises questions about ownership, copyright, access inequality, and the long-term preservation of digital files.

What does ‘metadata’ mean in digital archives?

Metadata is data that describes other data — in digital archives, it is structured information about a digital object that makes it discoverable, understandable, and usable. Metadata for a digitised manuscript might include title, author, date, language, provenance, and rights statement. Standards such as Dublin Core, MARC, and TEI provide common frameworks that enable interoperability between different digital collections. Poor metadata makes digital objects effectively invisible to researchers, so quality and consistency are crucial.

What is a digital corpus in humanities research?

A corpus (plural: corpora) in digital humanities is a structured collection of texts assembled for research purposes. Digital corpora enable computational analysis impossible with physical texts — searching for word patterns across thousands of documents, tracing the evolution of language over time, or comparing the style of different authors. Famous digital corpora include the British National Corpus and the HathiTrust Digital Library. Corpus design decisions — what texts to include and how to annotate them — fundamentally shape the research questions that can be asked.

What is topic modelling in digital humanities?

Topic modelling is a machine learning technique used to discover abstract themes within a large collection of documents. The most widely used algorithm is Latent Dirichlet Allocation (LDA), which identifies groups of words that tend to appear together. Historians have used it to trace the emergence and decline of political themes in newspapers, while literary scholars have used it to study genre conventions. The results require careful human interpretation, as the computer identifies statistical patterns rather than understanding meaning.

What is open access in digital scholarship?

Open access refers to the free, unrestricted online availability of research outputs — including articles, datasets, and digital editions — so that anyone can read, download, and reuse them without financial or legal barriers. The open access movement argues that publicly funded research should be freely available to the public. In digital humanities, open access is particularly important because DH projects often produce digital editions, datasets, and tools that need to be shared widely. Preprint servers, institutional repositories, and open-access journals are the main vehicles.

What is a digital edition in humanities?

A digital edition is a scholarly edition of a historical text that uses digital technologies to offer features not possible in print, such as multiple versions displayed in parallel, interactive annotations, hyperlinks to related documents, image facsimiles alongside transcriptions, and encoded markup enabling computational searching. The Text Encoding Initiative (TEI) provides XML-based guidelines for marking up digital editions. Famous examples include the Walt Whitman Archive and the Samuel Beckett Digital Manuscript Project.

What is GIS in digital humanities research?

GIS (Geographic Information Systems) is a technology for capturing, storing, analysing, and visualising spatial data. In digital humanities it is used to map historical phenomena such as the spread of diseases, trade networks, migration patterns, and battle locations. Historical GIS projects layer digitised maps with data about past events to reveal spatial patterns that text-based analysis cannot easily show. The Mapping the Republic of Letters project used GIS to visualise correspondence networks of early modern scholars across Europe.

How can I improve my digital humanities vocabulary in English?

Group terms by domain: archival (digitisation, metadata, finding aid, provenance), computational (text mining, corpus, algorithm, visualisation), publication (open access, digital edition, TEI, repository), and theory (distant reading, digital preservation). Reading the Digital Humanities Quarterly (open access) and the Programming Historian (free tutorials) exposes you to DH vocabulary in both theoretical and practical contexts. Joining DH online communities and following the #DH hashtag connects you with researchers using this vocabulary in real discussions. Trying beginner-level text analysis tools such as Voyant Tools helps you internalise computational vocabulary through hands-on practice.