Showing posts with label Reading Response LIS 501. Show all posts
Showing posts with label Reading Response LIS 501. Show all posts

Friday, November 20, 2009

Models and Standards for Authority Control

Models and standards work well when agreement is reached on which models and which standards to use for an activity or process. An additional requirement for enabling models and standards to meet the objectives of their design is that those working within the models should also adhere to the standards. This paper will synthesize the concepts found in the text with the comments from two information researchers about metadata quality and the problems of not adhering to standards.
In recent weeks, our class discussion focused on the description of information resources using metadata that allows for retrieval of data related to names, titles and subjects. Further, in order for relevant, or the best, information resources to be retrieved, the metadata and the retrieval system have to be subject to the same description process and authority control. As the vast volumes of information resources expanded, it became necessary to develop models and standards for this authority control.
Over time, superior models have been adopted and eventually accepted by the professional community of information scientists, managers and users. So that today, we have the general bibliographic models of Functional Requirements for Bibliographic Records (FRBR) and the Functional Requirements for Authority Data (FRAD). These two models outline the way that resource titles, their associated names, and associated subjects are related to one another and to other resources and also establish the way that those relationships are organized, categorized and classified. Like the models, the standards have also experienced lengthy development over time by way of constant improvement and upgrading until today several general standards exist. They are: Anglo-American Cataloguing Rules, Second Edition, 2002 Revision (AACR2); Dublin Core Agents, Metadata Authority Description Schema (MADS) and the soon to be launched Resource Description and Access (RDA) and the Statement of International Cataloguing Principles. These models and standards have been established, among other reasons, in order to improve metadata quality.
Many studies (cited in the Park and Salo articles) have shown that following these models and standards make retrieval of information, especially about bibliographic resources, possible and more accurate than if there were no standard methods of organization and classification. When the models and standards are followed, metadata quality is enhanced and information seekers have greater success in their information searches as more relevant information and greater amounts of information are found.
Examples of information sources that have been lacking in metadata quality and thus have led to ineffective searching for users are digital repositories and institutional repositories. Jung-ran Park identifies a number of studies on metadata quality that identify principles of “good metadata” and several other studies that help to identify the problems with metadata in digital repositories.
According to the National Information Standards Organization (NISO) there are six principles of good metadata, which are found on pg. 215 of the Park article. These principles are heavily related to standards and authority control. Obviously, following a standard makes the retrieval of data more successful because the standard has been intelligently established over time and will lead to better results in a search than if the data is organized, classified and indexed in a haphazard way. The essential key is to get submitters of data and the catalogers of data to follow the prescribed standards consistently and avoid making errors.
As mentioned, Park cites studies that uncover the problems with metadata quality in repositories. In summation of the many different problem criteria cited by the studies, Park simplifies our thinking on the subject by consolidating them into three common criteria: completeness, accuracy and consistency. She explains that by completeness she means that the metadata is complete enough to facilitate its purpose of making the resource to which it refers fully accessible in a search. Accuracy means correct entry of metadata and avoidance of mismatched metadata that is imported from data providers. She calls for consistency to be measured on two levels, the conceptual or semantic level and looking at the data format on the structural level.
Dorothea Salo, in her article “Name Authority Control in Institutional Repositories”, highlights these same problem criteria by showing how the methodology designed for the institutional repository did not protect against them. She may not use the same words, but she is emphasizing throughout her comments the problems with incomplete, inaccurate and inconsistent metadata.
According to Salo, the factors causing the problem in institutional repositories include the process of self-archiving by submitters of resources, ingesting material possessing no metadata control, and harvesting materials from aggregators and e-indexers which possess inconsistent metadata.
Obviously, Park’s recommendations for the improvement of repositories are to establish methods to improve the completeness, accuracy and consistency of metadata. Her recommendations include establishing metadata guidelines as a best practice for information organizations. This suggestion also applies to mechanized or automatic generation of metadata, leading into the further development of the semantic web technology. She cites several projects currently underway towards this objective. While Salo mentions several barriers to solutions to these problems with metadata inconsistencies, she does recommend finding a method to pass corrected metadata around between generators and managers of metadata. (I might add even involving users.) Salo suggests that use of the Open Archives Initiative and JISC are two paths to pursue vigorously.
Finding solutions to these problems is the challenge for the current generation of information scientists. Launching RDA and developing the Semantic Web will no doubt be the main focus for the near term. Establishing guidelines as best practices is one way, but this only gets results when others follow along. Development of the semantic web provides hope on the horizon, but again, information professionals engaged in the organization, cataloging and indexing of information resources must individually contribute to the completeness, accuracy and consistency of metadata in every aspect of their jobs in order for models and standards to take hold and take full effect. It will take personal commitment and engagement of all concerned.
Resources:
Taylor & Joudrey, Chapter 8. Metadata: Access and authority control (pp. 245-301).
Park, J.-R. (2009). Metadata quality in digital repositories: A survey of the current state of the art. Cataloging & Classification Quarterly 47(3/4): 213-228.
Salo, D. (2009). Name authority control in institutional repositories. Cataloging & Classification Quarterly 47(3/4): 249-261.

Friday, October 30, 2009

About Aboutness

Introduction
Subject analysis is a vastly important and challenging area of information organization. The literature reviewed in this week's reading assignments has not only outlined the process of subject analysis but brought to the fore several difficult aspects that have kept information professionals unsatisfied with its current condition in general, and with indexing for information retrieval systems specifically. According to Taylor & Joudrey (T&J)(pg. 303), being able to identify “precisely an item’s subject matter (often referred to as aboutness . . .)” is of significant value. Of equal value, and generally less definitively accomplished is then carefully and accurately assigning appropriate terms in an index to represent that aboutness. In this reading response, a brief synthesis is provided of the M. J. Bates (not so brief) article and the Taylor & Joudrey text. They are supportive of one another in their discussions of aboutness and the difficulties in achieving a database and retrieval system design that is completely satisfactory to its users. The conclusion is that there have been many changes in the subject analysis space since Bates wrote her article, and there is yet more on the near horizon that will further solve some of the inherent challenges.

Synthesis of Concepts
T&J informs that much literature is extant that points out that aboutness is more than just the words in the text used by an author. Accurate aboutness takes into consideration, among other things, the intentions of both the author and the users of the resource. This suggests there must be a connection between the person developing the vocabulary of the index and the end user, or at least an understanding of the latter by the former. Otherwise, the index is of great value to the indexer, but less so to the user. Bates (pg. 4) reinforces this concept by pointing out that the information seeking user doesn’t know what the resource is yet. The user has no idea what it is about, or what it can answer, or even how to ask for it. The indexer doing the analysis has the record in hand. There is no gap in understanding the resource. The main challenge for the indexer is to add to that “hands-on” knowledge by anticipating how the user will search for the resource. Even to the point of anticipating which terms the searcher might use, based on the information need. If the user’s guesses in the search don’t match the indexer’s description, then the resource is not found in the search and it is of little value to the user.

Further, Bates contrasts the indexer’s experience (with the subject analysis process) to the information seeking experience of the user. The indexer is highly trained and educated in the subject analysis arena, whereas the user is not, and may not even be experienced in the search process. So, having their thought process match is unlikely, unless the indexer thinks in terms of the potential searcher.
The main points of the Bates article are that human, database and domain factors impact the subject analysis and search processes. Those impacts can be negative if indexers don’t follow careful and well thought-out methodologies. Inherently, the indexer wishes to build something of great intellectual and scholarly value. Potentially, the more logical and sophisticated the retrieval system design is, the less it will resemble the inexperienced searching techniques of the user. The indexer must take into account the user community as well as the domain or environment of the users.

T&J describes the process of analysis and assigning terms. The text gives three examples of approaches (or methodologies) to the process, which are identified by their authors, Langridge, Wilson and Lancaster (as a representative of the Use-based Approach). Each approach consists of steps that generally follow a basic outline of asking questions. These questions begin in a sort of general arena and move to more specific characteristics of the information resource being analyzed. For instance, “What is it?”; “What is it for?”; and “What is it about?” in Langridge’s approach. Wilson describes in his approach four methods, which similarly work from generalities to specifics. First (Purposive), “What was the author’s purpose?”; second (Figure-Ground), “Is there a central figure that stands out?”; third (Objective), counting references to determine which vastly outnumber the others; and fourth (Cohesion), “what holds the work together?”. This general to specific question-tree approach was likened by Bates (pg. 14) to the same way library users approach the reference desk. They start with a general opening question and, based on the librarian’s interviewing skills, the interview moves to specifics until the librarian has determined what the user really wants and provides a path to that information. It seems that the database searcher follows this same approach of trying general terms and getting more specific until the search has been narrowed to relevant and desired information.

The impetus for following methodologies such as those presented in the text and article under review here is to overcome the many challenges inherent in communication of information. People from different environments, cultures and background simply think of things differently. Anticipating those differences will aid in overcoming them. The suggestion by Bates to build a user-friendly front-end to attach to the logically developed database makes good sense and since the time of her writing, this suggestion has been implemented by most developers of retrieval systems.

In today’s Web 2.0 environment, retrieval systems are advancing to the stages of incorporating user generated aboutness, or tagging, which will take the process a step further in overcoming some of the challenges identified in the Bates article and the T&J text. According to Rolla (pg. 178), user tagging such as is present in LibraryThing has added richness beyond the library catalog. This is the positive result that Bates, in particular, and the T&J chapter suggest we must pursue in the information organization profession as we overcome inherent challenges of subject analysis.

Conclusion
Many studies have been conducted and much has been written about the subject analysis process (see Notes in T&J, Chapter 9, pp. 328-332) which has added to the discussion before and since the 1998 Bates article. These studies and the literature have helped to formulate new methodologies and to enhance existing approaches to subject analysis. As approaches and methodologies, like those considered in this week’s reading material, continue to develop with an eye on solving the aboutness challenges discussed, the search process will only improve.

References:

Taylor & Joudrey, Chapter 9, Subject analysis (pp. 303-332) and Appendix A. An approach to subject analysis (pp. 419-427)

Bates, M.J. (1998). Indexing and access for digital libraries and the Internet: human, database, and domain factors. Journal of the American Society for Information Science 49(13): 1185-1205.

Rolla, P. J. (2009). Can user-supplied data improve subject access to library collections? Library Resources & Technical Services 53(3): 174-184.