Friday, November 20, 2009

Models and Standards for Authority Control

Models and standards work well when agreement is reached on which models and which standards to use for an activity or process. An additional requirement for enabling models and standards to meet the objectives of their design is that those working within the models should also adhere to the standards. This paper will synthesize the concepts found in the text with the comments from two information researchers about metadata quality and the problems of not adhering to standards.
In recent weeks, our class discussion focused on the description of information resources using metadata that allows for retrieval of data related to names, titles and subjects. Further, in order for relevant, or the best, information resources to be retrieved, the metadata and the retrieval system have to be subject to the same description process and authority control. As the vast volumes of information resources expanded, it became necessary to develop models and standards for this authority control.
Over time, superior models have been adopted and eventually accepted by the professional community of information scientists, managers and users. So that today, we have the general bibliographic models of Functional Requirements for Bibliographic Records (FRBR) and the Functional Requirements for Authority Data (FRAD). These two models outline the way that resource titles, their associated names, and associated subjects are related to one another and to other resources and also establish the way that those relationships are organized, categorized and classified. Like the models, the standards have also experienced lengthy development over time by way of constant improvement and upgrading until today several general standards exist. They are: Anglo-American Cataloguing Rules, Second Edition, 2002 Revision (AACR2); Dublin Core Agents, Metadata Authority Description Schema (MADS) and the soon to be launched Resource Description and Access (RDA) and the Statement of International Cataloguing Principles. These models and standards have been established, among other reasons, in order to improve metadata quality.
Many studies (cited in the Park and Salo articles) have shown that following these models and standards make retrieval of information, especially about bibliographic resources, possible and more accurate than if there were no standard methods of organization and classification. When the models and standards are followed, metadata quality is enhanced and information seekers have greater success in their information searches as more relevant information and greater amounts of information are found.
Examples of information sources that have been lacking in metadata quality and thus have led to ineffective searching for users are digital repositories and institutional repositories. Jung-ran Park identifies a number of studies on metadata quality that identify principles of “good metadata” and several other studies that help to identify the problems with metadata in digital repositories.
According to the National Information Standards Organization (NISO) there are six principles of good metadata, which are found on pg. 215 of the Park article. These principles are heavily related to standards and authority control. Obviously, following a standard makes the retrieval of data more successful because the standard has been intelligently established over time and will lead to better results in a search than if the data is organized, classified and indexed in a haphazard way. The essential key is to get submitters of data and the catalogers of data to follow the prescribed standards consistently and avoid making errors.
As mentioned, Park cites studies that uncover the problems with metadata quality in repositories. In summation of the many different problem criteria cited by the studies, Park simplifies our thinking on the subject by consolidating them into three common criteria: completeness, accuracy and consistency. She explains that by completeness she means that the metadata is complete enough to facilitate its purpose of making the resource to which it refers fully accessible in a search. Accuracy means correct entry of metadata and avoidance of mismatched metadata that is imported from data providers. She calls for consistency to be measured on two levels, the conceptual or semantic level and looking at the data format on the structural level.
Dorothea Salo, in her article “Name Authority Control in Institutional Repositories”, highlights these same problem criteria by showing how the methodology designed for the institutional repository did not protect against them. She may not use the same words, but she is emphasizing throughout her comments the problems with incomplete, inaccurate and inconsistent metadata.
According to Salo, the factors causing the problem in institutional repositories include the process of self-archiving by submitters of resources, ingesting material possessing no metadata control, and harvesting materials from aggregators and e-indexers which possess inconsistent metadata.
Obviously, Park’s recommendations for the improvement of repositories are to establish methods to improve the completeness, accuracy and consistency of metadata. Her recommendations include establishing metadata guidelines as a best practice for information organizations. This suggestion also applies to mechanized or automatic generation of metadata, leading into the further development of the semantic web technology. She cites several projects currently underway towards this objective. While Salo mentions several barriers to solutions to these problems with metadata inconsistencies, she does recommend finding a method to pass corrected metadata around between generators and managers of metadata. (I might add even involving users.) Salo suggests that use of the Open Archives Initiative and JISC are two paths to pursue vigorously.
Finding solutions to these problems is the challenge for the current generation of information scientists. Launching RDA and developing the Semantic Web will no doubt be the main focus for the near term. Establishing guidelines as best practices is one way, but this only gets results when others follow along. Development of the semantic web provides hope on the horizon, but again, information professionals engaged in the organization, cataloging and indexing of information resources must individually contribute to the completeness, accuracy and consistency of metadata in every aspect of their jobs in order for models and standards to take hold and take full effect. It will take personal commitment and engagement of all concerned.
Resources:
Taylor & Joudrey, Chapter 8. Metadata: Access and authority control (pp. 245-301).
Park, J.-R. (2009). Metadata quality in digital repositories: A survey of the current state of the art. Cataloging & Classification Quarterly 47(3/4): 213-228.
Salo, D. (2009). Name authority control in institutional repositories. Cataloging & Classification Quarterly 47(3/4): 249-261.

No comments:

Post a Comment