Understanding Digitalization: A Comprehensive Glossary for Libraries, Archives, and Knowledge Institutions
By Pushparaj Subedi
Library and Information Science | Digital Libraries | Knowledge Management |
Digital Preservation
Introduction
Digital technology is transforming how societies create, organize, preserve, share, and use information. Libraries, archives, museums, universities, government institutions, and research organizations are increasingly adopting digital technologies to improve access to knowledge, enhance services, preserve cultural heritage, and support education and research.
However, the terminology associated with digitalization can sometimes be confusing. Terms such as digitization, digitalization, digital transformation, digital preservation, digital archiving, and digital library are often used interchangeably, even though they refer to different concepts and processes.
A clear understanding of these terms is essential for information professionals, researchers, policymakers, educators, students, and anyone involved in developing digital information services.
This article presents a comprehensive glossary of digitalization-related terminology, with particular attention to libraries, archives, information management, knowledge management, and digital preservation. It is intended to serve as a practical reference for professionals and learners working in the digital knowledge environment.
1. Understanding the Fundamental Concepts
The starting point for understanding digitalization is to distinguish between several closely related concepts.
Digitization is the process of converting analog materials into digital representations. For example, scanning a printed book, photographing a historical manuscript, or converting an audio cassette into a digital audio file are forms of digitization.
Digitalization refers to the use of digital technologies to improve or redesign existing processes, services, and workflows. A library that introduces online circulation, electronic resource management, and digital reference services is undertaking digitalization.
Digital transformation involves broader organizational change through the integration of digital technologies into institutional strategies, structures, services, and working practices. For example, transforming a traditional library into an integrated knowledge service that connects digital collections, research data, learning platforms, and AI-assisted discovery may form part of a digital transformation initiative.
Other important concepts include:
- Digital content: Information created, stored, or delivered in digital form.
- Digital object: An identifiable digital item, such as an electronic document, photograph, video, dataset, or digitized manuscript.
- Digital asset: A digital object that has informational, cultural, institutional, or operational value.
- Digital infrastructure: The hardware, software, networks, storage, and platforms that support digital services.
- Digital ecosystem: The interconnected network of people, organizations, technologies, standards, and services involved in creating and using digital information.
- Digital literacy: The ability to find, evaluate, create, and communicate information effectively using digital technologies.
- Digital inclusion: Ensuring that individuals and communities can access and benefit from digital technologies.
- Digital divide: Inequalities in access to, use of, and benefits from digital technologies.
- Digital governance: The policies, responsibilities, and controls used to manage digital resources, platforms, and services.
Understanding these distinctions helps institutions plan digital initiatives more effectively and avoid treating the mere conversion of physical materials into electronic files as a complete digital transformation.
2. Digitization Technologies and Processes
Digitization involves a combination of equipment, technical procedures, file management, and quality control. The following terms are particularly relevant to digitization projects.
|
Term |
Definition |
|
Scanning |
Capturing a physical document or image using a scanner or imaging device. |
|
Digital imaging |
Creating digital representations of physical objects or scenes. |
|
Resolution |
The level of image detail captured or represented. |
|
DPI |
Dots per inch, a measure commonly associated with scanning and printing resolution. |
|
PPI |
Pixels per inch, a measure of pixel density in digital images. |
|
Bit depth |
The number of bits used to represent tonal or color information. |
|
Color management |
Controlling color reproduction across imaging devices and workflows. |
|
Optical Character Recognition (OCR) |
Converting images of printed text into machine-readable text. |
|
Handwritten Text Recognition (HTR) |
Recognizing and transcribing handwriting from digital images. |
|
Image enhancement |
Improving image visibility, contrast, or readability. |
|
Batch processing |
Processing multiple files or records through a common workflow. |
|
Metadata capture |
Recording information describing a digital object and its characteristics. |
|
Quality assurance (QA) |
Establishing processes to ensure that digitization outputs meet defined requirements. |
|
Quality control (QC) |
Inspecting outputs to identify errors, defects, or inconsistencies. |
|
Checksum |
A calculated value used to detect unintended changes or corruption in a file. |
For example, digitizing a rare Nepali manuscript may require careful handling, suitable imaging equipment, appropriate lighting, high-quality image capture, descriptive metadata, quality checks, and secure storage of the resulting files.
OCR may improve searchability, but it does not replace the original page image. Historical documents, complex scripts, damaged pages, and older printing styles may require manual transcription and verification.
3. Digital File Formats and Technical Standards
Selecting suitable file formats is a fundamental part of digital collection management. Different formats serve different purposes, including preservation, access, editing, and distribution.
|
Term |
Definition and common use |
|
TIFF |
A high-quality image format often used for preservation masters. |
|
JPEG |
A compressed image format commonly used for access and distribution. |
|
JPEG 2000 |
An image format supporting advanced compression and different image quality requirements. |
|
PNG |
An image format supporting lossless compression. |
|
|
A format for presenting and exchanging electronic documents. |
|
PDF/A |
A family of PDF standards intended for long-term preservation of electronic documents. |
|
WAV |
An audio format commonly used for uncompressed or high-quality audio preservation. |
|
MP3 |
A compressed audio format commonly used for distribution. |
|
FLAC |
A lossless compressed audio format. |
|
MP4 |
A multimedia container commonly used for video and audio. |
|
XML |
A markup language for representing structured information. |
|
JSON |
A lightweight format for exchanging structured data. |
|
CSV |
A text-based format for tabular data. |
|
Unicode |
A standard for representing characters across languages and writing systems. |
|
UTF-8 |
A widely used Unicode character encoding. |
|
Lossless compression |
Compression that preserves all original information. |
|
Lossy compression |
Compression that reduces file size by discarding some information. |
|
Open format |
A format with publicly documented specifications that can support implementation by different systems. |
|
Proprietary format |
A format controlled by a particular organization or vendor. |
A sound digitization strategy should distinguish between preservation masters and access copies. Preservation masters are created and managed according to defined quality requirements, while access copies may be compressed or optimized for online delivery.
The appropriate format depends on the type of material, intended use, preservation risks, storage capacity, accessibility requirements, and institutional policy.
4. Digital Preservation and Digital Archiving
Digitization creates digital files, but digital preservation ensures that those files remain authentic, accessible, and usable over time.
Digital preservation is particularly important for national libraries, public archives, universities, museums, research institutions, and organizations responsible for safeguarding documentary heritage.
Key terms include:
- Digital preservation: The policies, strategies, and actions required to maintain digital information and its usability over time.
- Digital archiving: The systematic acquisition, management, preservation, and provision of access to digital records and collections.
- Preservation master: A high-quality digital representation retained for preservation purposes.
- Access copy: A derivative file prepared for convenient use or online delivery.
- Fixity: The property of a digital object remaining unchanged.
- Bit rot: A term describing gradual data corruption or deterioration that may make digital information unreadable.
- Format obsolescence: The risk that software or technical environments required to interpret a file may no longer be available.
- Migration: Moving digital content to a different format, system, or environment to maintain usability.
- Emulation: Recreating an earlier computing environment to access digital materials that depend on obsolete software or hardware.
- Replication: Maintaining copies of digital information in multiple locations or systems.
- Media refreshing: Copying information onto new storage media to reduce risks associated with aging or failing media.
- Authenticity: Confidence that a digital object is what it claims to be.
- Provenance: Information about the origin, history, ownership, or custody of an object.
- Digital continuity: Maintaining the ability to locate, interpret, access, and use digital information over time.
- Disaster recovery: Procedures for restoring data and systems after disruption or loss.
Several standards and frameworks support this work. The Open Archival Information System (OAIS) reference model, standardized as ISO 14721, provides a conceptual framework for digital preservation. PREMIS supports preservation metadata, while the ISO 16363 framework addresses the audit and certification of trustworthy digital repositories.
Digital preservation should not be understood as simply saving files to a server or making backup copies. It requires documented responsibilities, technical controls, metadata, monitoring, preservation planning, and periodic review.
5. Metadata and Knowledge Organization
Metadata is fundamental to digital libraries and archives because it enables resources to be discovered, identified, understood, managed, and preserved.
Metadata can be divided into several categories:
- Descriptive metadata: Information used to identify and discover a resource, such as its title, creator, subject, and publication date.
- Structural metadata: Information showing how different files or components of a digital object relate to one another.
- Administrative metadata: Information required to manage and administer digital resources.
- Technical metadata: Information about file formats, dimensions, software, and other technical characteristics.
- Preservation metadata: Information documenting the management and preservation of digital objects over time.
- Rights metadata: Information about copyright, permissions, licences, and access restrictions.
Important standards and concepts include:
|
Term |
Purpose |
|
Dublin Core |
A widely used set of metadata elements for describing information resources. |
|
MARC 21 |
A standard for representing and exchanging bibliographic and related data. |
|
RDA |
Resource Description and Access, a standard for describing and providing access to resources. |
|
BIBFRAME |
A linked-data framework designed to support bibliographic description. |
|
METS |
A standard for encoding descriptive, administrative, and structural metadata. |
|
MODS |
An XML-based schema for bibliographic description. |
|
EAD |
A standard for encoding archival descriptions and finding aids. |
|
PREMIS |
A data dictionary supporting digital preservation metadata. |
|
IIIF |
The International Image Interoperability Framework, supporting interoperable delivery and presentation of digital images and audiovisual resources. |
|
Authority control |
The standardization of names, subjects, and identifiers to support consistent discovery. |
|
Controlled vocabulary |
An authorized set of terms used consistently for indexing and retrieval. |
|
Taxonomy |
A structured classification of concepts or subjects. |
|
Thesaurus |
A controlled vocabulary that defines relationships among terms. |
|
Ontology |
A formal representation of concepts and relationships within a domain. |
|
Linked data |
A method of connecting structured information through identifiers and relationships. |
Effective metadata is especially important in multilingual environments. In Nepal, digital collections may include Nepali and other languages, multiple scripts, transliterated names, historical spellings, and culturally specific subject terminology.
National digital initiatives should therefore consider multilingual metadata, Unicode compliance, authority control, consistent subject access, and interoperability across participating institutions.
6. Digital Libraries and Information Management Systems
A digital library is more than a website containing downloadable documents. It is an organized collection of digital resources supported by systems and services for discovery, access, management, and, where appropriate, preservation.
The following concepts are commonly used in digital library development:
- Digital library: An organized collection of digital resources supported by discovery and access services.
- Institutional repository: A platform for collecting, managing, preserving, and providing access to an institution's scholarly outputs.
- National digital library: A national initiative or service supporting access to digital information, education, research, or cultural heritage.
- Union catalogue: A combined catalogue representing holdings from multiple libraries.
- National thesis portal: A centralized discovery or access service for theses and dissertations.
- Repository: A system that stores, organizes, manages, and provides access to digital objects.
- Integrated library system (ILS): Software supporting library functions such as cataloguing, circulation, and acquisitions.
- Content management system (CMS): Software used to create, organize, and publish digital content.
- Digital asset management (DAM): Systems and practices for organizing, managing, and distributing digital assets.
- Discovery layer: An interface that helps users search across library collections and electronic resources.
- Federated search: Searching across multiple independent information systems.
- Full-text search: Searching the textual content of documents rather than only their descriptive metadata.
- Application Programming Interface (API): A mechanism through which software systems exchange data or request services.
- Interoperability: The ability of systems to exchange information and use it meaningfully.
- OAI-PMH: The Open Archives Initiative Protocol for Metadata Harvesting, used to collect metadata from participating repositories.
- Open access: Online access to scholarly outputs without a price barrier, subject to applicable rights and conditions.
- Open Educational Resources (OER): Educational materials released under open licences or otherwise in the public domain.
Examples of software used in this environment include DSpace for repositories, Koha for integrated library management, and WordPress for website and content management. Their suitability depends on institutional requirements, technical capacity, available resources, and long-term sustainability.
A national digital library may connect multiple institutional repositories rather than requiring every institution to deposit all its content in a single centralized system. The architecture should be selected according to governance, ownership, technical capacity, preservation responsibilities, and user needs.
7. Data Management and Research Information
Digitalization generates and connects large volumes of information. Effective data management is therefore necessary for research, institutional administration, service delivery, and evidence-based policymaking.
Important terms include:
- Data: Recorded facts, observations, measurements, or representations.
- Database: An organized collection of data managed for retrieval and use.
- Database Management System (DBMS): Software used to create, maintain, and query databases.
- Structured data: Data organized according to a predefined schema.
- Semi-structured data: Data organized through markers or flexible structures without a rigid tabular format.
- Unstructured data: Information without a fixed tabular structure, such as many documents, photographs, and recordings.
- Data governance: Policies, responsibilities, and controls for appropriate data management and use.
- Data stewardship: Operational responsibility for data quality, documentation, access, and responsible use.
- Data cleaning: Detecting and correcting errors, inconsistencies, and duplicates.
- Data integration: Combining information from different sources into a coherent resource.
- Data warehouse: A repository designed to support analytical reporting using integrated data.
- Data lake: A storage environment that accommodates data in multiple formats.
- Data mining: Identifying patterns or relationships within datasets.
- Data visualization: Presenting information through charts, maps, graphs, and other visual forms.
- Research Data Management (RDM): Planning, documenting, storing, sharing, and preserving research data.
- FAIR principles: Principles stating that data should be Findable, Accessible, Interoperable, and Reusable.
- Data lifecycle: The stages through which data passes, from creation and use to preservation or disposal.
- Data provenance: Information about the origin of data and the processes through which it has passed.
For universities and research organizations, research data management is increasingly important. A research project may produce survey responses, interview transcripts, statistical datasets, photographs, audio recordings, and analytical files. These materials require appropriate documentation, secure storage, access arrangements, ethical safeguards, and preservation planning.
8. Cloud Computing and Cybersecurity
Cloud computing enables institutions to use computing resources, software, and storage through network-based services. It can support digital libraries, institutional repositories, online learning, and national information infrastructure.
Common terminology includes:
- Cloud computing: Delivery of computing resources and services over a network.
- Software as a Service (SaaS): Software delivered as an online service.
- Platform as a Service (PaaS): A managed environment for developing and deploying applications.
- Infrastructure as a Service (IaaS): On-demand computing, networking, and storage infrastructure.
- Public cloud: Cloud infrastructure offered to multiple customers.
- Private cloud: Cloud infrastructure dedicated to a single organization.
- Hybrid cloud: An environment combining public cloud services with private infrastructure.
- Hosting: Providing computing infrastructure for websites, applications, or data.
- Encryption: Transforming information so that it cannot be read without the necessary decryption capability.
- Authentication: Verifying a user's identity.
- Authorization: Determining which resources and actions an authenticated user is permitted to access.
- Multi-factor authentication (MFA): Verifying identity using multiple types of evidence.
- Access control: Rules and mechanisms governing access to digital resources.
- Backup: A separate copy of information maintained for recovery.
- Firewall: A security control that filters network traffic according to defined rules.
- Malware: Software designed to damage, disrupt, or gain unauthorized access to systems.
- Phishing: Deceptive communication intended to obtain sensitive information or induce unsafe actions.
- Ransomware: Malware that typically encrypts data or threatens its disclosure to demand payment.
- Cybersecurity: Practices and technologies for protecting systems, networks, and information.
Digital libraries must consider security alongside accessibility. A repository may provide public access to approved documents while restricting administrative functions, confidential records, unpublished research, and copyrighted materials.
Institutions should also establish backup policies, access controls, incident response procedures, disaster recovery arrangements, and clear responsibilities for security management.
9. Artificial Intelligence and Emerging Digital Technologies
Artificial intelligence is introducing new possibilities for information discovery, document processing, metadata generation, research support, and user services.
Important terms include:
- Artificial Intelligence (AI): Technologies designed to perform tasks associated with human cognitive capabilities.
- Machine Learning (ML): Methods that enable computer systems to learn patterns from data.
- Deep learning: Machine learning based on multilayer neural networks.
- Generative AI: AI systems that generate content, including text, images, audio, video, and code.
- Large Language Model (LLM): A model trained on large amounts of text to process and generate language.
- Natural Language Processing (NLP): Computational methods for analyzing and processing human language.
- Computer vision: Techniques for analyzing and interpreting visual information.
- AI-assisted OCR: OCR using AI or machine learning to improve text recognition.
- Automated classification: Assigning documents or other resources to categories automatically.
- Semantic search: Searching based on meaning and context rather than exact word matches alone.
- Knowledge graph: A representation of entities and their relationships.
- Vector database: A database designed to store and search vector representations of information.
- Retrieval-Augmented Generation (RAG): A method combining information retrieval with generative AI to produce responses grounded in retrieved material.
- Hallucination: AI-generated information that is incorrect, unsupported, or fabricated.
- Prompt engineering: Designing instructions and inputs to guide AI system outputs.
- Human-in-the-loop: A process in which humans review, guide, or approve automated outputs or decisions.
- Explainable AI: Techniques intended to make AI outputs or decisions more understandable.
- Algorithmic bias: Systematic unfairness arising from data, model design, or deployment.
- Responsible AI: Developing and using AI with attention to fairness, transparency, privacy, accountability, and safety.
Libraries and archives can use AI-assisted tools to improve OCR, generate draft metadata, classify collections, support multilingual discovery, and help users navigate large information collections.
However, AI outputs require appropriate verification. Incorrect transcriptions, fabricated references, biased classifications, and inaccurate descriptions can undermine the integrity of digital collections. Human review remains important, particularly for rare manuscripts, culturally significant materials, sensitive records, and authoritative bibliographic descriptions.
10. Digital Governance, Copyright, and Ethical Responsibilities
Digitalization is not only a technical undertaking. It also raises legal, organizational, social, and ethical questions.
Important terms include:
- Copyright: Legal protection for qualifying original works.
- Creative Commons: A family of licences that allows creators to grant specified reuse permissions.
- Open licence: A licence permitting defined uses and redistribution under stated conditions.
- Public domain: Works not protected by applicable copyright restrictions or whose relevant protection has expired or been waived.
- Digital Rights Management (DRM): Technologies and controls governing access to or use of digital content.
- Data protection: Legal and organizational safeguards for personal information.
- Informed consent: Voluntary agreement based on adequate information about a proposed activity or use of data.
- Data sovereignty: The principle that data is subject to the laws and governance of the relevant jurisdiction.
- Digital ethics: Ethical principles guiding the development and use of digital technologies.
- Accessibility: Designing digital systems and content so that people with diverse abilities can use them.
- Universal design: Designing products and environments to be usable by as many people as possible.
- Audit trail: A record of actions and changes within a system.
- Risk assessment: Identifying and evaluating threats to information, systems, and services.
- Retention schedule: A documented timetable specifying how long records are kept and when they may be disposed of or transferred.
- Legal deposit: A legal requirement to deposit specified publications or materials with designated institutions.
- Vendor lock-in: Dependence on a vendor that makes changing systems or services difficult.
- Exit strategy: A plan for transferring data and responsibilities when leaving a platform or service provider.
For public and academic institutions, these issues must be addressed before digitization and online publication begin. Digitizing a document does not automatically authorize its public distribution. Institutions should assess copyright, ownership, privacy, consent, access restrictions, and applicable legal requirements.
A responsible digitalization policy should also consider accessibility, language diversity, community participation, and the needs of people with limited connectivity or digital skills.
11. Essential Standards for Digital Libraries and Archives
Standards provide a common technical and organizational foundation for creating, exchanging, managing, and preserving digital information.
Some widely used standards and frameworks include:
|
Standard or framework |
Application |
|
ISO 14721 (OAIS) |
Reference model for an open archival information system and long-term digital preservation. |
|
ISO 16363 |
Audit and certification framework for trustworthy digital repositories. |
|
ISO 15489 |
Principles and concepts of records management. |
|
ISO 23081 |
Metadata principles and concepts for records management. |
|
ISO 19264 |
Image quality analysis for cultural heritage imaging. |
|
Dublin Core |
Descriptive metadata for information resources. |
|
MARC 21 |
Bibliographic and related data exchange. |
|
PREMIS |
Digital preservation metadata. |
|
METS |
Encoding and transmission of digital object metadata and structural information. |
|
IIIF |
Interoperable delivery and presentation of digital images and audiovisual resources. |
|
OAI-PMH |
Metadata harvesting between participating repositories. |
|
WARC |
Storage of web crawls and associated archival information. |
|
BagIt |
Packaging digital content and manifests for transfer and storage. |
|
WCAG 2.2 |
Web accessibility guidance. |
|
Unicode |
Consistent representation of characters across languages and writing systems. |
|
RDF |
Representing structured information as linked data. |
The appropriate standards depend on the purpose of the system, the nature of the collection, technical capacity, and applicable requirements. Standards should be incorporated into documented workflows rather than treated merely as a checklist of acronyms.
12. Digitalization in the Context of Nepal
Nepal has a diverse information and documentary heritage, including historical manuscripts, printed books, government records, newspapers, photographs, maps, oral histories, academic theses, research reports, and audiovisual materials.
Digitalization offers opportunities to improve access to these resources, connect institutions, support education and research, and strengthen the preservation of documentary heritage.
Several concepts are particularly relevant to developing a coordinated national knowledge infrastructure.
National Digital Library: A coordinated initiative to facilitate the discovery and use of digital learning, research, cultural, and documentary resources.
National Union Catalogue: A shared discovery service that enables users to search bibliographic records and identify holdings across participating libraries.
National Thesis Portal: A platform for discovering or accessing theses and dissertations produced by universities and research institutions.
Distributed repository architecture: An approach in which institutions retain responsibility for their collections while using common standards and interoperable systems to exchange information.
Metadata harvesting: The collection of metadata from participating repositories to support centralized discovery without necessarily transferring the underlying files.
Multilingual metadata: Descriptive information supporting discovery across different languages and scripts.
National digital preservation framework: A coordinated approach defining preservation standards, institutional responsibilities, technical requirements, and long-term sustainability arrangements.
Digital knowledge infrastructure: The combined policies, technologies, standards, institutions, skills, and services required to support the creation, preservation, discovery, and use of knowledge.
Developing national digital services requires more than selecting software or purchasing storage. It involves institutional coordination, clear governance, sustainable financing, skilled personnel, copyright management, technical interoperability, user-centred design, and long-term preservation planning.
A distributed approach may allow libraries and universities to maintain control of their own collections while contributing metadata and approved digital content to national discovery services. The design of such a system should reflect national priorities, institutional capacities, and the needs of users.
13. Building Professional Capacity in Digitalization
Digitalization is a multidisciplinary field. Effective professionals need a combination of technical knowledge, information management skills, preservation expertise, and an understanding of organizational and user needs.
A practical learning pathway can be organized into five areas.
First, digitization fundamentals: Scanning, image quality, color management, OCR, file formats, metadata capture, and quality assurance.
Second, digital preservation: OAIS, PREMIS, fixity checking, checksums, migration, authenticity, backup strategies, and preservation planning.
Third, digital library systems: Repository platforms, integrated library systems, metadata standards, APIs, OAI-PMH, interoperability, and discovery services.
Fourth, governance and sustainability: Copyright, privacy, accessibility, cybersecurity, risk management, budgeting, policy development, and institutional responsibilities.
Fifth, advanced digital knowledge services: AI-assisted OCR, semantic search, linked data, knowledge graphs, research data management, analytics, and responsible AI.
Professional development should combine theory with practical experience. Training programmes can include scanning and quality assessment, metadata creation, repository configuration, preservation planning, digital collection audits, and the development of standard operating procedures.
Institutions should also invest in staff development, peer learning, professional networks, and collaboration among libraries, archives, universities, government agencies, and technology specialists.
Conclusion
Digitalization is a continuing process that connects technology with information, people, institutions, and knowledge. Its success depends not only on converting physical materials into digital files but also on organizing information effectively, preserving digital resources, enabling discovery, ensuring responsible access, and maintaining services over time.
For librarians, archivists, information scientists, educators, researchers, and policymakers, understanding digitalization terminology provides a foundation for informed decision-making and professional collaboration.
The concepts presented in this glossary can support the planning of digital libraries, institutional repositories, national union catalogues, thesis portals, archival digitization programmes, and digital preservation initiatives.
As digital technologies continue to evolve, the professional responsibility remains clear: to ensure that information is discoverable, trustworthy, accessible, meaningful, and available to future generations.
Digitalization should ultimately be understood not simply as a technological change, but as an opportunity to strengthen the creation, preservation, sharing, and use of knowledge for the benefit of society.
About the Author
Pushparaj Subedi is a library and information science professional with experience in library management, knowledge management, digital information services, and institutional information systems. His professional interests include digital libraries, digital preservation, information policy, knowledge sharing, and the development of accessible and sustainable knowledge infrastructures.
Keywords: #Digitalization, #digitization, #digital transformation, #digital libraries, #digital preservation, #digital archiving, #metadata standards, #knowledge management, #information management, #artificial intelligence, #national digital library, #digital repository, #Nepal.
No comments:
Post a Comment