Skip to content

How Clean Is Our Catalogue? A Quality Audit of 400 Retro-Converted MARC 21 Records in a College Koha Catalogue

  • 12 slides
  • 14 viva questions
  • 6 modules
  • No code needed

@retro-conversion-marc21-quality-auditUpdated Oct 2026

Counting what went wrong in a retro-conversion — wrong fields, stray class numbers, duplicate records — and writing the checklist that stops it

B.Lib.I.Sc, Cataloguing & Classification · Final sem · Beginner · 12 weeks · Solo

More info
Level
Beginner · 12 weeks · Solo
Relevant for
All India
Common at
Bangalore University, University of Mumbai, University of Delhi
Syllabus
Library-science departments University LIS · Project / dissertation (M.Lib Sem 4 typically; B.Lib project report) · Final semester
Tech stack
  • Koha ILS (staff interface, reports)
  • MARC 21 Bibliographic
  • AACR2 / RDA
  • DDC 23
  • Sears List of Subject Headings
  • MarcEdit (record review)
  • Excel (random sampling, error tallies, pivot tables)
For educational purposes only

Unlock this project

Full PPT + speaker notes, the step-by-step method, READMEFIRST, instructions and all 14 viva answers.

One-time. No subscription, no auto-renew, no drama.

Project packs

Credits never expire and work on any project. Use one here, save the rest for your friend who “will pay you back”.

  1. Pinned

    1 min

    Overview

    Many Indian college libraries have recently moved to Koha and "retro-converted" their old catalogue: staff, trainees or outsourced data-entry operators keyed years of accession-register and card-catalogue entries into MARC 21 records in a short time. The result is a catalogue that works, but nobody knows how accurate it is. This B.Lib.I.Sc project audits that accuracy.

    The study is set in Vidyanagar Arts and Science College Library (fictional), whose retro-conversion into Koha was finished last year. A random sample of about 400 bibliographic records is drawn from the catalogue and each record is checked against the book in hand (or its title page image) and against the rules: AACR2 (and RDA where the library has moved to it) for description and access points, MARC 21 for field and indicator use, DDC 23 for the class number, and the Sears List for subject headings. Duplicate records are searched separately.

    Every error is coded into a simple error taxonomy — access points, title and statement of responsibility (245), edition (250), publication details (260/264), physical description (300), classification (082), subject headings (650), item data, and duplicates — and tallied in Excel. The report presents error rates by field and by type, identifies likely root causes by interviewing the staff who did the conversion, and ends with a practical quality-control checklist the library can use for all new cataloguing.

    Syllabus alignment

    Library-science departments · University LIS

    Project / dissertation (M.Lib Sem 4 typically; B.Lib project report) · Final semester

    Subjects this project applies
    • Knowledge Organisation: Cataloguing Theory (AACR2 / RDA)
    • Cataloguing Practice (MARC 21)
    • Knowledge Organisation: Classification Theory and Practice (DDC)
    • Library Automation (Koha)
    • Research Methodology & Statistics
    How it is evaluated

    See your department's project guidelines.

    1 min read · 14 viva questions

  2. 2 min

    Synopsis

    Abstract

    Retro-conversion turns a library's legacy catalogue into machine-readable records. Done quickly and without review, it spreads small errors — a title in the wrong subfield, a missing author, a wrong class number — through thousands of records, and these errors silently hide books from users of the OPAC. This project conducts a sample-based quality audit of a college library's retro-converted Koha catalogue, measures error rates by field and type, investigates causes, and proposes a quality-control checklist.

    Introduction

    A catalogue is only as good as its records. Ranganathan's "every book its reader" depends on each book being findable by author, title and subject. In automated systems, where searches depend on specific MARC fields and indexes, a mistake in the wrong field can make a book practically invisible. Yet most small libraries never audit their records after migration.

    Review of literature (gap)

    LIS literature on retro-conversion in India focuses mainly on methods (in-house keying versus outsourcing versus copy cataloguing) and on software. Studies that measure record quality tend to address large university catalogues or union catalogues. Small affiliated college libraries, where conversion is often done by temporary staff, are rarely studied. This project provides a simple, repeatable audit method suited to such libraries.

    Existing practice vs proposed practice

    • Existing: records entered from registers with minimal review; no written local cataloguing policy; no routine check of new records; duplicates noticed only by chance.
    • Proposed: periodic sample audits with a fixed error taxonomy; a short local cataloguing policy; a quality-control checklist applied to every new record; a monthly duplicate check using Koha reports.

    Feasibility

    • Technical: the audit needs only Koha's staff interface, MarcEdit for viewing records in bulk, and Excel.
    • Operational: checking about 400 records takes roughly three to four weeks for one student working a few hours a day.
    • Ethical: no personal user data is involved; staff interviews are anonymous and focus on processes, not individuals.
  3. 1 min

    Problem statement

    Vidyanagar Arts and Science College Library completed the retro-conversion of its catalogue into Koha in a short period using temporary staff and trainees working from accession registers and old catalogue cards. Since then, library staff and students have noticed books that exist on the shelf but cannot be found in the OPAC, the same title appearing as several records, and class numbers that place books in the wrong section. No one knows how widespread these problems are or which kinds of errors are most common.

    The problem this project addresses is: what is the level and pattern of errors in the library's retro-converted MARC 21 records, what causes them, and what practical quality-control measures can prevent them in future? The answer must be based on a systematic sample rather than anecdotes, so that the library can decide where to spend limited correction effort.

  4. 1 min

    Objectives & scope

    1. 01Draw a simple random sample of about 400 bibliographic records from the library's Koha catalogue.
    2. 02Develop an error taxonomy based on AACR2/RDA, MARC 21, DDC 23 and the Sears List of Subject Headings.
    3. 03Check each sampled record against the physical book and the rules, and code every error found.
    4. 04Estimate error rates by field, by error type and by severity (affects retrieval or cosmetic).
    5. 05Identify duplicate bibliographic records using Koha reports and title searches.
    6. 06Investigate likely causes through interviews with staff involved in the conversion.
    7. 07Propose a quality-control checklist and a short local cataloguing policy for the library.

    Scope

    In scope

    • Books (monographs) catalogued during the retro-conversion; one library.
    • Bibliographic fields 020, 082, 100/110, 245, 250, 260/264, 300, 490, 650 and the item field 952.
    • Duplicate record detection within the catalogue.
    • Staff interviews about the conversion process.

    Out of scope

    • Serials, theses, non-book material and e-resources.
    • Authority control files (noted as a recommendation only).
    • Correcting the entire catalogue; only the sample is corrected, and the method is handed over.
  5. 1 min

    Methodology

    The study uses a descriptive, sample-based audit design with a small qualitative component (staff interviews).

    PhaseWeeksWorkOutput
    1. Literature and rules1–2Review retro-conversion literature; list the rules for each field from AACR2/RDA, MARC 21, DDC 23 and SearsRule sheet, draft error taxonomy
    2. Sampling frame3Export the list of biblionumbers from a Koha report; generate random numbers in Excel; select about 400 recordsSample list
    3. Pilot4Check 30 records, refine the taxonomy and coding sheet, agree severity levels with the guideFinal coding sheet
    4. Record checking5–8Pull each book from the shelf (or use title-page photos), compare with its record, code every errorCompleted coding sheet
    5. Duplicate check8Koha reports on matching ISBNs and titles; manual confirmationDuplicate list
    6. Interviews9Short interviews with the librarian and two or three staff/trainees involvedInterview notes
    7. Analysis10–11Pivot tables: errors per record, error rate by field and type, severity; link to causesTables and charts
    8. Report12Findings, checklist, policy, report writingBound report

    Error measures: percentage of records with at least one error; mean errors per record; error rate for each field (records with an error in that field ÷ records where the field applies); share of errors that affect retrieval. Sample size logic: about 400 records gives a margin of error of roughly five percentage points at 95% confidence for a large catalogue — explain this in the report with the standard formula.

  6. 1 min

    Architecture & tech stack

    • Koha ILS (staff interface, reports)
    • MARC 21 Bibliographic
    • AACR2 / RDA
    • DDC 23
    • Sears List of Subject Headings
    • MarcEdit (record review)
    • Excel (random sampling, error tallies, pivot tables)

    The study design is summarised in the flowchart below.

    flowchart TD
      A[Rules: AACR2 / RDA, MARC 21, DDC 23, Sears] --> B[Draft error taxonomy]
      C[Koha report: list of biblionumbers] --> D[Random sample of about 400 records]
      B --> E[Pilot check of 30 records]
      D --> E
      E --> F[Final coding sheet and severity levels]
      F --> G[Check each record against the book]
      G --> H[Code errors in Excel]
      C --> I[Duplicate detection reports]
      I --> H
      J[Staff interviews] --> K[Root causes]
      H --> L[Error rates by field, type and severity]
      L --> K
      K --> M[Quality-control checklist and local cataloguing policy]

    Error taxonomy (coding sheet columns)

    CodeAreaExamples
    AAccess points (100/110/700)author missing, inverted wrongly, corporate body as personal name
    TTitle and responsibility (245)subtitle not in $b, wrong first indicator, statement of responsibility missing
    EEdition (250)edition in title field, "2nd ed." missing
    PPublication (260/264)place/publisher swapped, year missing or wrong
    DPhysical description (300)pages missing, illustrations not recorded
    CClassification (082 / 952$o)wrong DDC number, local number not matching 082
    SSubject (650)missing, not from Sears, too broad
    IItem data (952)barcode or item type wrong
    XDuplicate recordsame edition catalogued twice

    Each error is also marked R (affects retrieval) or C (cosmetic).

  7. 6 modules

    Modules

    • Rule Sheet & Error Taxonomy

      Summarises the AACR2/RDA, MARC 21, DDC 23 and Sears rules for each audited field and turns them into a coded error taxonomy with examples, refined after a 30-record pilot.

    • Sampling

      Exports all biblionumbers with a Koha SQL report, generates random numbers in Excel and selects about 400 records, documenting the procedure so another person could draw an equivalent sample.

    • Record Checking

      Compares each sampled record with the book's title page and verso, codes every error by area and severity on the coding sheet, and keeps notes and photos for unclear cases.

    • Duplicate Detection

      Runs Koha reports that group records by ISBN and normalised title, then checks candidates manually to confirm true duplicates of the same edition.

    • Root-Cause Interviews

      Conducts short, anonymous interviews with staff and trainees who performed the conversion about sources used, training, time pressure and review, and relates their answers to the error patterns.

    • Analysis & Checklist

      Builds pivot tables and charts of error rates by field, type and severity, and converts the most frequent errors into a one-page quality-control checklist and a short local cataloguing policy.

  8. Locked

    Presentation

    12 slides with speaker notes. The outline below is free; the bullets, notes and the generated .pptx unlock with the project.

    1. How Clean Is Our Catalogue?
    2. Background
    3. Objectives
    4. Rules used
    5. Sampling
    6. Error taxonomy
    7. Checking procedure
    8. Results: overall
    9. Results: by field
    10. Causes
    11. Quality-control checklist
    12. Conclusion

    Bullets, speaker notes and the .pptx download unlock with the project.

    Presentation is locked: 12 slides, Speaker notes, .pptx download.

  9. 7 sources

    References

    1. MARC 21 Format for Bibliographic Data — Library of Congress
    2. Koha Manual — Koha Community
    3. RDA Toolkit
    4. Dewey Decimal Classification and Relative Index, 23rd edition — OCLC
    5. Sears List of Subject Headings
    6. Gorman, M. — The Concise AACR2
    7. Ranganathan, S. R. — The Five Laws of Library Science

    Cite this bundle

    OnlyProjects. (2026). How Clean Is Our Catalogue? A Quality Audit of 400 Retro-Converted MARC 21 Records in a College Koha Catalogue: B.Lib.I.Sc Cataloguing & Classification project bundle [Educational resource]. https://onlyprojects.online/projects/blib-cataloguing-retro-conversion-marc21-quality-audit

Slides, diagrams & files

12 slides. Titles are free; bullets, speaker notes and the .pptx unlock with the project.

  1. SLIDE 1

    How Clean Is Our Catalogue?

  2. SLIDE 2

    Background

  3. SLIDE 3

    Objectives

  4. SLIDE 4

    Rules used

  5. SLIDE 5

    Sampling

  6. SLIDE 6

    Error taxonomy

  7. SLIDE 7

    Checking procedure

  8. SLIDE 8

    Results: overall

  9. SLIDE 9

    Results: by field

  10. SLIDE 10

    Causes

  11. SLIDE 11

    Quality-control checklist

  12. SLIDE 12

    Conclusion

Architecture diagram

1
flowchart TD
  A[Rules: AACR2 / RDA, MARC 21, DDC 23, Sears] --> B[Draft error taxonomy]
  C[Koha report: list of biblionumbers] --> D[Random sample of about 400 records]
  B --> E[Pilot check of 30 records]
  D --> E
  E --> F[Final coding sheet and severity levels]
  F --> G[Check each record against the book]
  G --> H[Code errors in Excel]
  C --> I[Duplicate detection reports]
  I --> H
  J[Staff interviews] --> K[Root causes]
  H --> L[Error rates by field, type and severity]
  L --> K
  K --> M[Quality-control checklist and local cataloguing policy]

Files

Viva questions & answers

3 of 14 questions free. Explain each answer in your own words before you move on.

  1. Concept

    What is retro-conversion?

    Retrospective conversion is the process of converting a library's existing manual catalogue records, from cards or registers, into machine-readable records in an automated system. It can be done by keying in-house, outsourcing or copying records from other catalogues.

  2. Concept

    What is the difference between AACR2 and RDA?

    AACR2 is a set of cataloguing rules organised around the type of material and the card-catalogue tradition. RDA is its successor, based on the FRBR and later library reference models, focusing on entities and relationships and on transcribing information as found. Many Indian libraries still follow AACR2 in practice.

  3. Concept

    Why does a wrong MARC field matter if the text is correct?

    Koha indexes specific fields and subfields for author, title and subject searches. If a title or author is placed in the wrong field, the text exists but is not indexed where the user searches, so the book becomes hard or impossible to find through the OPAC.

+11 more questions

They and the answers unlock with the project. Try answering the ones above yourself first. Your examiner will.

For educational purposes only. Use this bundle to understand how the project works, then build and write your own. Submitting it verbatim is between you, your conscience and your external examiner.