Epistemological Asymmetries in the Digital Archive: Cultural Bias in LLaMA‑1 and CommonCrawl‑Based Large Language Models and Its Implications for Digital Humanities Practice

Mayuree Pal ORCID ,  C.P. Rashmi ORCID
    Received: 11 February 2026; Revised: 1 May 2026; Accepted: 13 May 2026; Published: 25 June 2026

    Abstract

    Given the increasing integration of large language models (LLMs) in the digital humanities (DH) for textual analysis, archival research, and computational interpretation, an urgent need to decolonize LLM knowledge structures has emerged. This study conducts a secondary data analysis to highlight profound epistemological asymmetries embedded within LLaMA-1 and other documented CommonCrawl-based large language models and digital archives. The quantitative analysis reveals a stark representational disparity: approximately 82% of LLaMA-1's training set is explicitly English-filtered, and 49–56% of top-website content is in English, despite only 25.9% of global internet users speaking English as a first language. After more than a decade of systematic funding and infrastructure development, the digitisation of cultural heritage within European institutions remains stagnant at merely 16–23%. Crucially, no equivalent systematic surveys of cultural heritage digitisation exist for institutions in the Global South, indicating a severe global structural deficit. To mitigate these cascading disparities (from biased training data to uneven archival digitisation), this paper proposes the Epistemological Asymmetry Framework (EAF). The EAF serves as a structured, verifiable tool designed for DH practitioners and research institutes to audit, measure, and actively address cultural and linguistic biases. Ultimately, establishing equitable digital practices requires both rigorous multilingual corpus auditing and significant infrastructural investments in global heritage digitisation.

    Keywords

    References

      ×