Showing posts with label metadata. Show all posts
Showing posts with label metadata. Show all posts

Wednesday, May 7, 2008

David Weinberger: The New Shape of Knowledge

Keynote Speaker David Weinberger

“The New Shape of Knowledge” 11:00am-12:00pm

David Weinberger is something of a pioneer in understanding and explaining the internet. He is a frequent contributor to NPR’s “All Things Considered.” Here is a link to some of the stories on NPR http://www.npr.org/search.php?text=David+Weinberger

I also included a link to the Berkman Center for Internet and Society at Harvard University where David is currently a Fellow. Here you can find a bio and links to some of his works http://cyber.law.harvard.edu/people/dweinberger

I feel obligated to note that David was a comedy writer for Woody Allen, and he was also a humor columnist for a daily newspaper. Who ever thought librarians and humor don’t mix. I always thought humor and librarianship go well together, and my suspicions are being confirmed at this conference! Humor lovers unite!

David Weinberger is probably best known for the book he coauthored, “Cluetrain Manifesto,” which explained to business what the internet was really about. Check out this link with info and chapters from the book http://www.cluetrain.com/ Just keep scrolling down for further explanation, and a chance to read the entire book online for free.

David began discussing the New Yorker article on digitization and is discontents. We are in an age of abundance—an abundance of crap. We are getting pretty good at dealing with the amazing amount of crap available. There is also a large amount of good stuff that is available.


Order

Our natural inclination to manage and cluster objects and ideas is inherent in everyday life. We put spices together in the cabinet, pair fruits together on the counter, and meats in the fridge. We “lump and split” with our laundry—socks together, sweaters in one spot, etc. That is the way the world works—completely binary. We have this idea that there is a single way to organize things which is limiting thinking.


Entropy—A Philosophy of Change

This limited type of thinking is inherent in newspapers by privileging stories in a limited space on the front page. But in the age of information we are moving from the physical to another level of ordering known as digital. In time just about everything will be digitized. The stuff that is digitized it will 99% of the time be the first resort for information, including archives. This includes a rethinking of classification, known today as tagging. This is true through layers of linking. Just link to any blog and try to “follow the trail.” We have so many layers of metadata in our digital world that creates a huge mess; however, this mess, David argues, is a virtue.


Metadata

This is obvious in our search online at a library. Searching for Herman Melville will yield multiple data points, which is linked to a number of resources. The linking creates the metadata which accesses many data points. We no longer deal with one source of information, one piece of data, or a simple order. We are, David reminds us, moving from simple to complex order.

Unowned Order

The new order means that we get to manage the way we order information. We don’t need to make decisions for others based on our ordering. Instead, we can include everything. This can be overwhelming, but it gives people more options. We can manage the incredible amount of information by tagging. There will no doubt be other ways to manage the information.

Library of Congress--flickr

This reordering is making its way into the now fledgling(?) Library of Congress. LOC put tons of photos on flickr with some metadata—not a lot, but enough. There was an enormous amount of tagging and posts that discussed the organization of the photos being shared online. Comments and tagging contribute to the reorganization of material on the web. This also becomes a learning environment. People posting ask questions and engage others in what they see. The dialogue becomes a rather abstract exercise in epistemology, or what we know, and how we know what we know. This rationalization and questioning is on display in the interactive approach online. There is a re-evaluative process online.

Wikipedia

This re-evaluation process is evident in Wikipedia where contributors share ideas and open discussion about information that is posted through comments and challenges in acknowledging human fallibility, reorganizing notions of dialogue. We don’t see posts disputing the verity of an article in the New York Times or the Boston Globe. Fallibility is exposed in Wikipedia.

Reconsidering Knowledge

Knowledge, or notions of knowledge, is refashioned in the “mess” of the information age. Our knowledge is not necessarily changing, but the process is. Maybe we are in an age of process, or an age of transition. Whatever it is, it sure beats the Dark Ages (so I hear from someone who was there), and I am certainly enjoying the process whatever happens with knowledge. And David says it will take at least a generation to work through this new knowledge.

Next Generation Data Format - Eric Lease Morgan - University of Nortre Dame - May 7, 2008

Eric Lease Morgan started with "its not about 'what we do', it is about 'how we do it'."

How we got to where we are

Morgan reminded us that we have been using services (LC cards, OCLC, etc.) to help us build catalogs for generations and using the MARC record since 1965. It is a sequential specific data structure. Then, the invasion of Dialogue, CD-ROMS, network connections, and jump ahead to today, where Information structures, such as Google, are based on different indexing principles.

Redefine the definition of librarianship

Traditional core mission of librarianship is to collect, organize, archive, and disseminate data, information, and knowledge for our respective communities.

Here's what Morgan suggests:

Reduce dependence

Librarians should have more control. They should find ways to use technology to provide services. Using Open Source Software is one way.

Exploit technology

Morgan advocates using relational databases, indexing systems (Googlebox, Keynote Search, etc.), XML.

If we want to communicate to other places, in other words beyond the library community, we must speak other languages, and XML is the language being used.

As a "step in the right direction", conversion programs can help change MARC to MARC XML. The idea of describing a work is valid, but the method does not work in the broader world. The "Secret Codebook", better known as MARC, is limited because you need the code to interpret the fields.

A word from the audience (Dodie Gaudet) mentioned that the tags in XML are now in English, but MARC using numbers is more universal.

Commercial products use indexers to create interface. For example: Koha uses Zebra; Aqua developed a proprietary Indexer.

Work collaboratively

Work with peers and stakeholders in the inside and outside the library world.

"Next generation" library catalogs

Provide service not just catalog systems. As a part of the community, use this base to create services at the local level. Add, annotate, cite, compare and contrast, plot a map, find similar, trace authors, etc.

In response to a question, Morgan suggested see the applications (Scriblio, WordPress, VuFind, etc.) demonstrated during the pre-conference (MLA May 6th).

Thursday, May 3, 2007

Future of MARC -- Dr. Bill Moen

Future of MARC: the Challenges and Opportunities of 21st Century Cataloging
[Eventually, the link to this presentation will be on the MLA Conference website]

  • "We need to be adding value through our practices....We'd better be saving people time and money."
  • "We need to be meeting the needs of our users.
  • Less focus on the methods. Our methods should be invisible and unobstructive. How can we take our structures and hide them, but not hide the power that they provide?

There's nothing wrong with thinking in terms of market share -- that's our reality. We no longer have a lock on the target market that we used to. We have a limited set of resources that are availble through our catalog, and now users can see how limited they are compared with everything else that's out there.

What do we mean when we say "MARC"?
Record format, as defined by ISO 2709/ANSIZ39.2, & structural elements of the format -- this is going to go away and be replaced by markup languages
Metadata scheme -- defined by MARC 21 and fields, subfields, indicators and their semantics

Approaching MARC's future:

Requirements for a record format/metadata scheme
Goldsmith & Knudson's Requirements:
Granularity -- how fine a detail can you get to
Transparency --
Extensibility --

Roy Tennant's Requirements (slide went by too fast)

McCallum's 10 Format Attributes
  • XML
  • Granularity
  • Versatility
  • Extensibility
  • Modularity
  • Hierarchy support
  • Crosswalks
  • Tools
  • Cooperative Management
  • Pervasive


Functional Requirements for Bibliographic Records (FRBR)
Produce a framework that would provide a clear, precisely stated and commonly shared understanding of what it is that the bib record aims to provide information about and what it is that we expect the record to achieve in terms of answering user needs.
Based on Entity-Relationship modeling (work, expression, manifestation, item / persons, corporate bodies)
An important part of the FRBR report was the focus on users & user tasks: find, identify, select, obtain
If bib records are not supporting user tasks, what's the point?

FRBR is introducing new vocabulary and a new understanding of the items we catalog. We're seeing the implementation of these ideas in new library catalogs.
We currently ask our users to put up with a lot of noise, lots of individual record, rather than letting them actively drill down with limited choices

Responding to recent developments
No AARC3; Resource Description and Access (RDA), which are more guidelines on content creation, and have separation from syntax or record format

Library Systems and Data Formats (wiki) -- grassroots efforts to look at

Metadata
  • Essential in library applications
  • Variety of metadata schemes
  • Variety of functions and services supported
  • Increasing use of machine-generated metadata - there aren't enough catalogers in the world
  • Role of handcrafted metadata needs continuing review & assessment


These are not threats to the livelihood of catalogers/TS librarians. There's plenty of work around, but they have to change the approach to handcrafted metadata -- where's the value added of that hands-on work.

Looking at empirical data
The cataloging record you create is an artifact that reflects decisions, policies and choices, and can be investigated to see patterns and needs.
Catalogers create metadata that can be very rich (MARC).

There had never been a study before of exactly how catalogs actually construct MARC records and what they actually do. So, Dr. Moen did one.
There is a *lot* of redundancy in the records. The 80/20 rule holds: 4% of fields/subfields accounted for 80% of occurrences, 96% of all fields accounted for 20% occurrences (where occurrence = data in the field)

MCDU Project -- Reports containing results of analysis of utilization, commonly used elements
OCLC gave them the entire WorldCat as of May 2005, so they're working from the whole dataset
82 hours for a script to process and load the records as MySQL and 258 GB = mad data set
Millions of book records
Categories of Questions: General profile of the dataset & actual numbers of occurrences
167 fields used
14 fields accounted for 80% of all occurrences
21 fields accounted for 90% of all occurrences
110 fields occur in less than 1% of all records
"656 a:" occurred in 1 record of out 7.5 million. Why?

They also looked at field/subfield combinations. They keep finding a small core of elements that are commonly used -- are these what catalogers should be focusing on? Is this what the machine-generated cataloging should be focusing on?

Making Sense of Numbers
Not interpreting the value of an individual fields, but looking at patterns and larger recommendations/guidelines. Comparing the data to FRBR user tasks.Is there a common core of elements that are used? Is there a threshold below which things just aren't used so much?

Are library catalogers providing data to support FRBR tasks?
In MCDU dataset, only 59 fields/subfields (13% of total) occur at or above the threshold of use in OCLC book records.

Questions for consideration
  • What is needed in a bib record? Are catalogers working too hard and creating stuff no one uses?
  • Support for four user tasks? In the context of FRBR, what does it mean to support a user task?
  • How can we use metadata for effective management of information resources?
  • How do your systems use the infrequently used data? What about the 62% of all fields used in less than 1% of records?
  • Can we argue persuasively for the cost/benefit for your existing practice?
  • Should the focus be on the high-value, high-impact, high-quality data in a few fields/subfields? Can we identify these? What would it mean to costs of cataloging to focus this way? What would this mean for training new catalogers?
  • Can MCDU results inform your local practices?
  • What metadata scheme will we use? It won't be MARC.
  • [missed one]


Confluence of change -- all of this data and the realities of life now mean that change will happen. It's just a question of when and will we be prepared?

Is the next study a look at which fields users are using? Is this useful to know? What data about how users search is useful? There have been a lot of studies done of how users search in the Search Engine world; how do these translate into the library field?

[Jenn sez: There are a lot of stunned heads in this audience. As a non-cataloger, I'm hoping this hasn't just caused heart attacks throughout the room.]

"Focus on the needs that they meet, not the methods that they use."

That quote from Karen Calhoun, in her presentation On Competition for Catalogs and Catalogers.

Some notes from her talk this morning:

Rethink the catalog in light of a changed world: the current model is broken.
  • Users not getting what they want
  • Content has changed, users have changed
  • Library service model must change
  • Catalog must change, and cataloging must change
Disintermediation: decrease in guided access to content; users are more self-sufficient
How do we save users' time in this new arena, when the service model is not longer library-centric?
Shift in user preferences for Web-based info & multimedia formats

Popularity of digital resources as primary sources (e.g. Valley of the Shadow, http://valley.vcdh.virginia.edu/)

Social networks can fulfill information needs (e.g. LinkedIn, http://www.linkedin.com/)

"Catalogers have been through hell in the past ten years."

A new kind of cataloger:
  • examines assumptions
  • involved w/ all types of info objects
  • moves to next gen systems & services
  • make info more visible and easier to use
  • metadata and beyond
Focus on the needs of particular disciplines & communities
Example: http://vivo.library.cornell.edu - both a resource portal (through a curated index) and social networking site for life sciences at Cornell
(Blogger's note: this looks like a great resource!)

Opportunities for cataloging and catalogers:
  • More digitization --> more full-text search --> more metadata. Metadata recycling & reuse will be crucial.
  • Digitization projects are an opportunity for catalogers
  • Archives and special collections are on the rise; there's likely to be "a ton of work" for interested catalogers

Increasing visibility of collections and services
  • We need to be where our users' eyes are; their eyes aren't on library webpages
  • Offsite storage as a challenge to browsing
  • Partnerships
  • Robust, interconnected discovery & content delivery systems

Metadata is a strategic issue, yet 'library-type' metadata will need to be reexamined

Blurring of lines between public services & technical services

For more on this topic, check out the forthcoming piece in Library Hi Tech - Being a Librarian: Metadata and Metadata Specialists in the Twenty-first Century: http://dspace.library.cornell.edu/handle/1813/2231

Karen's report to the LOC: www.loc.gov/catdir/calhoun-report-final.pdf

On competittion for catalogs and catalogers

Karen Calhoun - how cool it is that MLA Tech. Services got Karen here to speak. I went to the NETSL annual conpsyched ference in April and her report was a main topic of conversation. I'm really psyched to be able to hear her talk. I'll try to do it justice here

Karen's presentation will be on the MLA web site after the conference - go look at it - it's been great!

Report name: Changing nature of the catalog and its integration with other discovery tools

http://www.loc.gov/catdir/calhoun-report-final.pdf

Report has been controversial. KC will talk about why
Users are anot getting what they need from online libraries and catalogs
contect has changed
users have changed
library service model must change
catalog itself MUST change
implication - Cataloging MUST change

KC is doing a review of how we used to work - very divided among types of jobs. No longer relevant.

Being a 21st century librarian:
Technology driven research, teaching and learning
Disinntermediation (decrease in guided access to content)
Global 'infosphere'
Accelerating shift in information seekers' preferences for web-based information and multimedia formats.

Our current service mode is Geocentric - local catalog is the sun. New model is Heliocentric - local catalog is a planet. New is a use centered view rather than a library centered view.

Social networks - competition to catalogs. Next widespread network is "linked in" for business communities. These are not compatible with library catalogs. in april there were 5.8 million second life users. second life library and info island are 2 library related worlds within second life.

what we think of as cataloging has to change. Catalogers must change, too.
KC says that catalogers are the most flexible folks in the library profession - downsizing, local systems, constant rebuilding of serials checkins, batcholads, macros, etc. metadata, moving to offsite storage. Continue through all this to crank out the books and othermaterials that users demand.
Are we ready for the NEXT phase???????

A new kind of cataloger:
Examines assumptions
Be involved with information objects of all types
Move to next generation systems and services
Make info (inc. but not limted to library collections) more visible and easier to use
Metadata and beyond

Talking about Vivo at Cornel. http://vivo.library.cornell.edu/
Metadata pushed to the limit and even further

2% of students begin their information searches on their library web sites! the other 98% start with a web browser. Is this a crisis or an opportunity for librarians?


digitization projects are becomming more and more important - archival collections are becomming more important. Many libarians will begin to spend money on these. Catalogers need to learn about archives to reposnd to this.

A new way to work - Instead of being a hoarder of contations, the library muyst become the facilitator of retrieval and dissemination - William Wulf, 2003

Save the time of the reader - 4th law of Ranganathan, 1931.
Metadata - Cooperative cataloging - great time saver but geting more and more expensive. Metadata is a strategic issue for libraries. Library type metadata needs to be changed and this cannot be ignored.

Please see slides 24, 25 on the MLA web site - lots of info - too much to get down here.

Catalogers have accomplished much in the past successfully, but there is a danger of myopia in
the profession now. The catalog is not the only valid way to describe collections.

Slides 26, 27 talk about the future of the job of 'cataloging'. May not be called catalogers anymore - maybe metadata specialists. Lots of bluring of lines between public and technical services. Lots of team work that cross traditional lines. Great need for info technology knowledge.

See Library Hi Tech, v. 25, no2. Preprint 17 December 2004
Being a librarian: metadata and metadata specialists in the twenty-first century.

http://dspace.library.cornell.edu/handle/1813/2231

Karen's remarks today taken from this article.

Questions -
Is VIVO a public site?
Answer - Yes


Interesting that only 2% of students start with the library web site. However, it has been observed that 95% of librarians begin with a search engine, too.

Answer Karen agrees - important to look in the 'personalities' why people search how they search.

Question about metadata standards and what other kinds of skills and knowledge to future catalogers need.

Answer Reuse of metadata - stuf that comes from outside the library that we reuse - need to know how to do that. Workflow analysis and quality improvement (see slide 17 on the right hand side for an example) . People who can look at the big picture

Question about controlled headings - where is it going
LCSH - we can do better. If we can create tools that would link with LCSH and create an algorithmic approach that would be more creative. Move our approach away from one record at a time. Cataloging theory is becomming more relevant now that our 'tools' are expanding. Tradtional forms need to connect with new forms.

Comment - please get rid of 'Cookery'!

Question - What can Karen suggest as to how to train people who are working in cataloging right now to use metadata.

Answer - there are classes to be taken, but if you do that and come back and not use it of course you forget it. See article - Digital Collections on a shoestring. Try to find something in your library that you can digitize. try to figure out how to do something!

Comment: The mindset of metadata specialists - want to do the absolute best job - too often using blinkers and not looking at the process from beginning to end. Start at the beginning and work up to the 'handcrafted' best. Think of the process as a continuum to build up your skills. This is often a team based project - metadata specialst, librarian, IT.

Comments - we still are going to have to run the back end of things - ILL, patron links, all that functionality of the ILS, the behind the scenes management back ends. How will we link up to allthese back end workings?

This was a GREAT program. Thanks to MLA Tech.Services Section for getting it together.